发布日期:2026-09-12
收录条目:20
1. Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
- 来源:MarkTechPost
- 发布时间:2026-09-11 22:01 UTC
- 链接:https://www.marktechpost.com/2026/09/11/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize/
摘要:ByteDance Seed, SUTD, Georgia Tech, M-A-P, and TokenWave.AI introduce HarnessDev, a benchmark that scores the runnable harness a model builds rather than the answer it returns. Starting from a seed that scores 0, 6 creat
2. Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
- 来源:MarkTechPost
- 发布时间:2026-09-11 21:05 UTC
- 链接:https://www.marktechpost.com/2026/09/11/anthropic-adds-plugin-evals-to-claude-code-6-grader-types-a-no-plugin-baseline-and-a-ci-gate-for-skills/
摘要:Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugi
3. Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- 来源:The Verge AI
- 发布时间:2026-09-11 20:44 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/994207/chatgpt-new-mexico-lawyer-fined-murder-appeal
摘要:New Mexico's Supreme Court is punishing a lawyer for including AI-fabricated witnesses and fake police testimony in an appeal for his client's murder conviction, according to a report from Reuters. In a filing on Wednesd
4. Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations
- 来源:AWS ML Blog
- 发布时间:2026-09-11 18:26 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/monitoring-production-agent-lifecycle-with-aws-devops-agent-and-agentcore-evaluations/
摘要:Multi-agent systems fail in ways traditional monitoring misses. This post presents a dual-layer approach to monitoring production agents: Amazon Bedrock AgentCore Evaluations for continuous quality scoring and AWS DevOps
5. Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload
- 来源:AWS ML Blog
- 发布时间:2026-09-11 18:24 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/beyond-the-price-per-token-choosing-the-right-openai-model-on-amazon-bedrock-for-your-workload/
摘要:Comparing models on dollars per million tokens misses what production workloads actually pay for: outcomes. This post shares an open-source benchmarking harness that measures cost per correct answer, agent trajectory cos
6. Build interactive MCP Apps using Amazon Bedrock AgentCore
- 来源:AWS ML Blog
- 发布时间:2026-09-11 18:23 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/build-interactive-mcp-apps-using-amazon-bedrock-agentcore/
摘要:Learn how to build and deploy an MCP App with interactive HTML widgets on Amazon Bedrock AgentCore. Because MCP Apps is a host-agnostic standard, the same server delivers the same rich experience across AI hosts like Cha
7. Anthropic spent this week in hot water over cybersecurity
- 来源:The Verge AI
- 发布时间:2026-09-11 16:09 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/994064/anthropic-spent-this-week-in-hot-water-over-cybersecurity
摘要:After admitting earlier this year that its AI models had hacked other companies' systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents dis
8. Meta says it’s changing AI suggestions after posing invasive personal questions
- 来源:The Verge AI
- 发布时间:2026-09-11 14:25 UTC
- 链接:https://www.theverge.com/tech/993974/meta-ai-prompt-invasive-suggestions
摘要:Meta says it's making changes to the prompts suggested by its AI chatbot after a viral video showed it digging for personal information about a woman's young daughters, as reported earlier by Futurism. In a statement to
9. Rapidly scaling online storage to serve over 1 billion ChatGPT users
- 来源:OpenAI News
- 发布时间:2026-09-11 10:00 UTC
- 链接:https://openai.com/index/scaling-storage-one-billion-users-part-one
摘要:Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.
10. Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
- 来源:MarkTechPost
- 发布时间:2026-09-11 06:57 UTC
- 链接:https://www.marktechpost.com/2026/09/10/cohere-releases-north-small-translate-a-218b-moe-translation-model-that-scores-83-6-on-wmt26-across-50-languages/
摘要:Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token and scores 83.6 on Cohere's WMT26 evaluat
11. Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
- 来源:MarkTechPost
- 发布时间:2026-09-11 06:34 UTC
- 链接:https://www.marktechpost.com/2026/09/10/sakana-ai-launches-fugu-max-and-fugu-ultra-v2-for-cheaper-stronger-multi-agent-orchestration/
摘要:Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M toke
12. Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation
- 来源:MarkTechPost
- 发布时间:2026-09-11 06:01 UTC
- 链接:https://www.marktechpost.com/2026/09/10/google-research-releases-toolgrad-answer-first-framework-hits-99-8-pass-rate-for-tool-use-data-generation/
摘要:Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then writes the matching user query. Guided by textual "gradients" fro
13. OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09203
摘要:arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outp
14. Adaptive Entangled Game Modules in Artificial General Intelligence
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09226
摘要:arXiv:2609.09226v1 Announce Type: new Abstract: We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral
15. Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09233
摘要:arXiv:2609.09233v1 Announce Type: new Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusab
16. Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09306
摘要:arXiv:2609.09306v1 Announce Type: new Abstract: This paper investigates the hypothesis that the first-order structure of physical interactions, i.e. gradients or Jacobians, characterizes the structure of phenomenal exper
17. An Autonomous GeoAI Agent for Arctic Eco-Navigation
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09374
摘要:arXiv:2609.09374v1 Announce Type: new Abstract: Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand seasonal accessibility while simultaneously introducing substantial oper
18. The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09395
摘要:arXiv:2609.09395v1 Announce Type: new Abstract: Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of availa
19. Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09413
摘要:arXiv:2609.09413v1 Announce Type: new Abstract: Choosing a recovery process for scale-up requires connecting laboratory results with product requirements, process costs, and scale effects. We analyze records from Pacific
20. Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
- 来源:arXiv cs.AI
- 发布时间:2026-09-11 04:00 UTC
- 链接:https://arxiv.org/abs/2609.09418
摘要:arXiv:2609.09418v1 Announce Type: new Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly ad