Administrator
发布于 2026-09-13 / 1 阅读
0
0

AI 每日资讯 - 2026-09-13

发布日期:2026-09-13

收录条目:20

1. Perplexity trusts GPT-6 Astra with end-to-end systems

摘要:Perplexity uses Astra to write communications, change software, and monitor production systems, and checks in much less frequently than with earlier models.

2. Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

摘要:Cognition, the company behind the Devin coding agent, has released SWE-2, its most capable coding model to date. SWE-2 is post-trained with reinforcement learning from Kimi K3, Moonshot AI’s 2.8T-parameter open model. Co

3. OpenAI’s rogue AI tried to hack another company in May

摘要:In May, hundreds of malicious and spam packages were uploaded to RubyGems, causing a serious disruption for the host. Now independent researchers have said that a swarm of OpenAI agents were responsible for the attack. N

4. Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’

摘要:OpenAI CEO Sam Altman confirmed that there would be no OpenAI IPO in 2026 during an interview with Fortune. Over the course of 45 minutes, Altman discussed a variety of subjects including the Hugging Face hacking inciden

5. Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

摘要:The Fly Language Model (FLM) drives all 166,700 retained neurons and 25.6 million edges of the MaleCNS fruit fly connectome with token embeddings, then adds a small learned correction to a frozen LFM2.5-1.2B-Instruct bac

6. Anthropic CEO says it’s time to pump the brakes on AI

摘要:Anthropic CEO Dario Amodei says the time has come to slow down AI development and will give third-party evaluators like METR access to its models to help ensure its "adherence to safety practices and commitments." In a w

7. Trump is giving data centers a pass to pollute

摘要:President Donald Trump is weakening environmental regulations in the name of speeding up the construction of AI data centers, raising health risks for Americans, a cadre of former EPA officials said this week in a briefi

8. OpenAI just wants to win

摘要:OpenAI has spent the last few years planting flags across the increasingly difficult terrain in mathematics. This week, it claimed one of its biggest prizes yet: a solution to a legendary Millennium Prize problem. In nor

9. Probabilistic Focal Search: Accelerating Bounded-Suboptimal Search via Lower-Bound Advancement

摘要:arXiv:2609.10584v1 Announce Type: new Abstract: Bounded-suboptimal search seeks a solution within a factor $w$ of optimal while reducing search effort. Focal Search (FS) uses heuristic guidance within FOCAL, the frontier

10. Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

摘要:arXiv:2609.10629v1 Announce Type: new Abstract: Quadratic Unconstrained Binary Optimization (QUBO) is a central formulation for combinatorial optimization and has gained increasing attention due to its compatibility with

11. A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

摘要:arXiv:2609.10654v1 Announce Type: new Abstract: The Abstraction and Reasoning Corpus (ARC) benchmarks cognitive generalization, the ability to infer and apply abstract rules from limited examples. This paper presents a m

12. Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning

摘要:arXiv:2609.10656v1 Announce Type: new Abstract: Selecting LoRA rank for diffusion fine-tuning requires balancing quality and compute cost. We present a controlled study on CIFAR-10 using a DDPM U-Net with ranks {2,4,8,16

13. Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

摘要:arXiv:2609.10657v1 Announce Type: new Abstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{w

14. An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

摘要:arXiv:2609.10712v1 Announce Type: new Abstract: We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we t

15. Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

摘要:arXiv:2609.10724v1 Announce Type: new Abstract: Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and depe

16. Towards a Deterministic Math Solver for Clinical Language Models

摘要:arXiv:2609.10728v1 Announce Type: new Abstract: Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard resp

17. Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

摘要:arXiv:2609.10824v1 Announce Type: new Abstract: Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural

18. When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents

摘要:arXiv:2609.10873v1 Announce Type: new Abstract: Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. We argue that update admission must be assessed through both error cont

19. Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows

摘要:arXiv:2609.10964v1 Announce Type: new Abstract: Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but al

20. Demystifying the Privacy-Utility Trade-off in LLM Interactions

摘要:arXiv:2609.10992v1 Announce Type: new Abstract: The integration of Large Language Models into daily tasks relies on context-rich instructions, inevitably exposing sensitive user information. Current privacy-preserving me


评论