发布日期:2026-08-04
收录条目:20
1. Evaluating Multimodal Vision Models with Moonshot PerceptionBench Using Robust Data Loading and Automated Judging
- 来源:MarkTechPost
- 发布时间:2026-08-03 22:26 UTC
- 链接:https://www.marktechpost.com/2026/08/03/evaluating-multimodal-vision-models-with-moonshot-perceptionbench-using-robust-data-loading-and-automated-judging/
摘要:In this tutorial, we design an end-to-end evaluation workflow for PerceptionBench. This multimodal benchmark measures fine-grained visual perception capabilities across tasks such as OCR, counting, localization, contextu
2. How to Secure AI Agents, MCP Servers, and LLM Apps in Production
- 来源:MarkTechPost
- 发布时间:2026-08-03 20:16 UTC
- 链接:https://www.marktechpost.com/2026/08/03/how-to-secure-ai-agents-mcp-servers-and-llm-apps-in-production/
摘要:AI agents, MCP servers, and LLM apps break the core AppSec assumption that applications do what their code says. This guide walks through a practical see-fix-protect framework: a five-layer agentic AI attack surface map,
3. Europe’s AI labeling and transparency rules are now in effect
- 来源:The Verge AI
- 发布时间:2026-08-03 17:38 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/974571/eu-ai-act-transparency-labels-rules-deepfakes
摘要:The European Union has ushered in some additional rules that aim to make it easier for people to identify chatbots and AI deepfakes online. The new transparency obligations under the bloc's landmark AI Act came into effe
4. From weeks to minutes: How Formula 1® uses agentic AI on AWS to accelerate data operations
- 来源:AWS ML Blog
- 发布时间:2026-08-03 17:24 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/from-weeks-to-minutes-how-formula-1-uses-agentic-ai-on-aws-to-accelerate-data-operations/
摘要:Formula 1® partnered with AWS to build the Data Accelerator, using agentic AI on Amazon Bedrock AgentCore to transform its MarTech data platform. Learn how F1 cut data source onboarding from up to 8 weeks to about 40 min
5. Automated Reasoning policy refinement in Amazon Bedrock
- 来源:AWS ML Blog
- 发布时间:2026-08-03 16:30 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/automated-reasoning-policy-refinement-in-amazon-bedrock/
摘要:Amazon Bedrock now supports automatic Automated Reasoning policy refinement. The refinement engine diagnoses failing tests and proposes formal-logic fixes for rule issues and language issues, and you approve every change
6. China’s Alibaba takes another swipe at America’s AI supremacy
- 来源:The Verge AI
- 发布时间:2026-08-03 11:01 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/974342/alibaba-qwen-max-open-weight-ai
摘要:Chinese tech giant Alibaba released what it says is its largest and "most capable AI model to date," claiming performance rivaling the best systems from US frontier labs Anthropic and OpenAI, as well as domestic rivals l
7. Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to Date
- 来源:MarkTechPost
- 发布时间:2026-08-03 08:24 UTC
- 链接:https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/
摘要:Alibaba's Qwen team moved Qwen3.8-Max from preview to general availability, with published per-token pricing and open weights due next week. The 2.4T parameter MoE model accepts text, image and video input across a 1M-to
8. Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths
- 来源:MarkTechPost
- 发布时间:2026-08-03 07:28 UTC
- 链接:https://www.marktechpost.com/2026/08/03/ogent-ai-team-releases-vr-1/
摘要:Cogent AI team released Cogent VR-1, a reasoning model post-trained specifically for cybersecurity rather than picking up cyber capability as a side effect of general coding strength. It ships with two companions: Intrus
9. How we built a realtime system for responsive voice AI in six months
- 来源:OpenAI News
- 发布时间:2026-08-03 07:00 UTC
- 链接:https://openai.com/index/continuous-voice-interaction-with-gpt-live
摘要:GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
10. OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28629
摘要:arXiv:2607.28629v1 Announce Type: new Abstract: The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agenti
11. Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28631
摘要:arXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI
12. LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28632
摘要:arXiv:2607.28632v1 Announce Type: new Abstract: Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation of conjectures with substantial ma
13. ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28642
摘要:arXiv:2607.28642v1 Announce Type: new Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue th
14. TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28657
摘要:arXiv:2607.28657v1 Announce Type: new Abstract: Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the ch
15. Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28659
摘要:arXiv:2607.28659v1 Announce Type: new Abstract: Cross-domain sequential recommendation (CDSR) aims to model users' dynamic interest transitions and sequential patterns across multiple domains. Recently, generative recomm
16. An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28662
摘要:arXiv:2607.28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same pers
17. How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28674
摘要:arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods re
18. Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28677
摘要:arXiv:2607.28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for sym
19. ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28678
摘要:arXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning.
20. Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO
- 来源:arXiv cs.AI
- 发布时间:2026-08-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.28679
摘要:arXiv:2607.28679v1 Announce Type: new Abstract: Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial inspection in factories. A particula