发布日期:2026-07-04
收录条目:20
1. Mistral AI Releases Leanstral 1.5: An Apache-2.0 Lean 4 Code Agent Model Solving 587 of 672 PutnamBench Problems
- 来源:MarkTechPost
- 发布时间:2026-07-03 22:20 UTC
- 链接:https://www.marktechpost.com/2026/07/03/mistral-ai-releases-leanstral-1-5-an-apache-2-0-lean-4-code-agent-model-solving-587-of-672-putnambench-problems/
摘要:Mistral AI released Leanstral 1.5, a free Apache-2.0 code agent model for Lean 4. It saturates miniF2F and solves 587 of 672 PutnamBench problems. The 119B mixture-of-experts activates 6.5B parameters per token. We break
2. Designing a Schema-Guided Invoice Intelligence Pipeline with lift-pdf for Accounts-Payable Extraction, Validation, and Ledger Generation
- 来源:MarkTechPost
- 发布时间:2026-07-03 21:25 UTC
- 链接:https://www.marktechpost.com/2026/07/03/schema-guided-invoice-intelligence-pipeline-with-lift-pdf/
摘要:In this tutorial, we build an end-to-end accounts-payable extraction pipeline with lift-pdf, using synthetic invoice PDFs as controlled test documents and a structured JSON schema as the target output format. Instead of
3. Anthropic wants to develop its own drugs
- 来源:The Verge AI
- 发布时间:2026-07-03 13:56 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/961311/anthropic-claude-science-ai-drug-development
摘要:At the event "The Briefing: AI for Science" earlier this week, Anthropic announced Claude Science, a new "AI workbench for scientists" that pulls fragmented tools and datasets into one environment, and generates figures
4. A behind-the-scenes look at Midjourney’s medical scanner leaves many questions unanswered
- 来源:The Verge AI
- 发布时间:2026-07-03 11:49 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/961265/midjourney-medical-ultrasound-scanner-behind-the-scenes-video
摘要:Midjourney has shown more of its futuristic medical scanner. It still hasn't shown much proof it works. The AI startup, best known for generating images, released a behind-the-scenes video of its dunk-tank ultrasound sca
5. Meet WebBrain: An Open-Source, Local-First AI Browser Agent That Reads Pages and Automates Tasks in Chrome and Firefox
- 来源:MarkTechPost
- 发布时间:2026-07-03 05:55 UTC
- 链接:https://www.marktechpost.com/2026/07/02/meet-webbrain-an-open-source-local-first-ai-browser-agent-that-reads-pages-and-automates-tasks-in-chrome-and-firefox/
摘要:WebBrain is a free, MIT-licensed AI browser agent for Chrome and Firefox. It reads pages, extracts data, and automates multi-step tasks through Ask and Act modes. Run it on local models like llama.cpp or Ollama for priva
6. PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01306
摘要:arXiv:2607.01306v1 Announce Type: new Abstract: Counterfactual explanations explain machine learning predictions by identifying minimal input changes that would alter a model's decision. Although many existing methods su
7. Auto-FL-Research: Agentic Search for Federated Learning Algorithms
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01366
摘要:arXiv:2607.01366v1 Announce Type: new Abstract: Federated learning (FL) research often depends on many small but consequential algorithmic choices: optimizer variants, server aggregation rules, local training schedules,
8. The Wiola Architecture for Efficient Small Language Models
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01394
摘要:arXiv:2607.01394v1 Announce Type: new Abstract: We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with any existing model family includi
9. Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01425
摘要:arXiv:2607.01425v1 Announce Type: new Abstract: Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Existing code summarizati
10. When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01426
摘要:arXiv:2607.01426v1 Announce Type: new Abstract: Autonomous customer-service agents are shifting from conversational interfaces toward operational execution roles: they retrieve firm records, apply service policies, and e
11. CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01433
摘要:arXiv:2607.01433v1 Announce Type: new Abstract: Divergent thinking is a crucial aspect of creativity, yet large language models (LLMs) tend to consistently generate similar responses to open-ended questions, in what has
12. Discrete Diffusion Language Models for Interactive Radiology Report Drafting
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01436
摘要:arXiv:2607.01436v1 Announce Type: new Abstract: Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competitive with autoregres
13. Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01465
摘要:arXiv:2607.01465v1 Announce Type: new Abstract: Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means hitting the right e
14. World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01470
摘要:arXiv:2607.01470v1 Announce Type: new Abstract: Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for RL from world feedb
15. Procedural Memory Distillation: Online Reflection for Self-Improving Language Models
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01480
摘要:arXiv:2607.01480v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the p
16. The Agentic Garden of Forking Paths
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01507
摘要:arXiv:2607.01507v1 Announce Type: new Abstract: Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are di
17. Janus: a Playground for User-Involved Agentic Permission Management
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01510
摘要:arXiv:2607.01510v1 Announce Type: new Abstract: AI agents that autonomously execute tool calls on a user's behalf raise pressing questions about permission management: what role could users play, and what role should the
18. Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01511
摘要:arXiv:2607.01511v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent reasoning capabilities in large language models. However, most existing CoT meth
19. OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01531
摘要:arXiv:2607.01531v1 Announce Type: new Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networks are flexible but
20. Scaling Trends for Lie Detector Oversight in Preference Learning
- 来源:arXiv cs.AI
- 发布时间:2026-07-03 04:00 UTC
- 链接:https://arxiv.org/abs/2607.01567
摘要:arXiv:2607.01567v1 Announce Type: new Abstract: Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses l