发布日期:2026-06-27
收录条目:20
1. Building Supervised Fine-Tuning Data from NVIDIA Open-SWE-Traces: Trajectory Parsing, Patch Analysis, Token Budgets, and Tool-Use Metrics
- 来源:MarkTechPost
- 发布时间:2026-06-27 00:02 UTC
- 链接:https://www.marktechpost.com/2026/06/26/building-supervised-fine-tuning-data-from-nvidia-open-swe-traces-trajectory-parsing-patch-analysis-token-budgets-and-tool-use-metrics/
摘要:In this tutorial, we work with NVIDIA's Open-SWE-Traces dataset to study agentic software-engineering trajectories for fine-tuning. We stream the data directly from Hugging Face, so we can process it efficiently in Googl
2. Cursor Study Finds Reward Hacking Inflates Coding-Agent Benchmark Scores on SWE-bench Pro
- 来源:MarkTechPost
- 发布时间:2026-06-26 23:31 UTC
- 链接:https://www.marktechpost.com/2026/06/26/cursor-study-finds-reward-hacking-inflates-coding-agent-benchmark-scores-on-swe-bench-pro/
摘要:A Cursor study shows coding agents retrieve known fixes instead of deriving them, inflating SWE-bench Pro scores through runtime contamination. The post Cursor Study Finds Reward Hacking Inflates Coding-Agent Benchmark S
3. Perplexity Launches Computer for Counsel: A Multi-Model Agentic Layer for Legal Workflows
- 来源:MarkTechPost
- 发布时间:2026-06-26 19:31 UTC
- 链接:https://www.marktechpost.com/2026/06/26/perplexity-launches-computer-for-counsel-a-multi-model-agentic-layer-for-legal-workflows/
摘要:Perplexity's Computer for Counsel extends Perplexity Computer to legal teams. It routes 20+ models across Midpage, MCP connectors, and Microsoft 365, with cited outputs lawyers can verify. The post Perplexity Launches Co
4. OpenAI Previews GPT-5.6 With Sol, Terra, and Luna: Tiered Models, New Reasoning Modes, Limited Access
- 来源:MarkTechPost
- 发布时间:2026-06-26 19:18 UTC
- 链接:https://www.marktechpost.com/2026/06/26/openai-previews-gpt-5-6-with-sol-terra-and-luna-tiered-models-new-reasoning-modes-limited-access/
摘要:OpenAI's GPT-5.6 family adds tiered models with max and ultra reasoning. Here is what early-level engineers should know. The post OpenAI Previews GPT-5.6 With Sol, Terra, and Luna: Tiered Models, New Reasoning Modes, Lim
5. OpenAI unveils GPT-5.6 amid US AI regulatory drama
- 来源:The Verge AI
- 发布时间:2026-06-26 17:00 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/957845/openai-gpt-5-6-trump-administration-ai-preview
摘要:Less than 24 hours after news broke that OpenAI would stagger its next model release at the request of the Trump administration, that model, GPT-5.6, is here. On Friday, the company unveiled the limited preview of its ne
6. Build interactive PDF text extraction from Amazon S3
- 来源:AWS ML Blog
- 发布时间:2026-06-26 14:47 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/build-interactive-pdf-text-extraction-from-amazon-s3/
摘要:In this post, you’ll build a server that extracts text from PDF files in Amazon S3 in real time. This protocol-based approach provides programmatic document access. You’ll walk through the architecture, set up the server
7. How Cara pioneers domain-specific AI for enterprise insurance brokerages with AWS
- 来源:AWS ML Blog
- 发布时间:2026-06-26 14:42 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/how-cara-pioneers-domain-specific-ai-for-enterprise-insurance-brokerages-with-aws/
摘要:In this post, we explore how Cara, built in cooperation with AWS, addresses these challenges. We walk through the technical design decisions and the AWS services that support the solution. We also share measurable outcom
8. Production-grade AI agents for financial compliance: Lessons from Stripe
- 来源:AWS ML Blog
- 发布时间:2026-06-26 14:38 UTC
- 链接:https://aws.amazon.com/blogs/machine-learning/production-grade-ai-agents-for-financial-compliance-lessons-from-stripe/
摘要:In this post, you learn how Stripe built a production-grade AI agent system for financial compliance. We cover the technical architecture of Stripe’s ReAct agent framework and the infrastructure decisions behind a dedica
9. Anthropic’s Mythos mess is only getting worse
- 来源:The Verge AI
- 发布时间:2026-06-26 14:07 UTC
- 链接:https://www.theverge.com/ai-artificial-intelligence/957327/anthropic-mythos-fable-ai-trump-administration-negotiations
摘要:It's been two weeks since Anthropic took its Mythos-class models offline after a Friday evening ultimatum from the Trump administration. The company sprang into action immediately, sending a barrage of executives to Wash
10. Previewing GPT-5.6 Sol: a next-generation model
- 来源:OpenAI News
- 发布时间:2026-06-26 10:00 UTC
- 链接:https://openai.com/index/previewing-gpt-5-6-sol
摘要:OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most advanced safety stack.
11. Meet container: Apple’s Open-Source Swift Tool for Running Linux Containers as Lightweight VMs on Apple Silicon
- 来源:MarkTechPost
- 发布时间:2026-06-26 08:48 UTC
- 链接:https://www.marktechpost.com/2026/06/26/meet-container-apples-open-source-swift-tool-for-running-linux-containers-as-lightweight-vms-on-apple-silicon/
摘要:Apple released container 1.0, an open-source Swift tool running Linux containers as lightweight virtual machines on Apple silicon. The post Meet container: Apple’s Open-Source Swift Tool for Running Linux Containers as L
12. Build a Nanobot-Style AI Agent in Google Colab with Tool Calling, Session Memory, Skills, and MCP Servers
- 来源:MarkTechPost
- 发布时间:2026-06-26 08:00 UTC
- 链接:https://www.marktechpost.com/2026/06/26/build-a-nanobot-style-ai-agent-in-google-colab-with-tool-calling-session-memory-skills-and-mcp-servers/
摘要:In this tutorial, we build a lightweight personal AI agent inspired by the architecture of nanobot, runnable entirely in Google Colab. We start from a provider abstraction, then add tool registration, session memory, lif
13. Detecting and Controlling Sycophancy with Cascading Linear Features
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26155
摘要:arXiv:2606.26155v1 Announce Type: new Abstract: Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavi
14. Life After Benchmark Saturation: A Case Study of CORE-Bench
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26158
摘要:arXiv:2606.26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the oppor
15. Refusal Lives Downstream of Persona in Chat Models
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26161
摘要:arXiv:2606.26161v1 Announce Type: new Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate m
16. AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26173
摘要:arXiv:2606.26173v1 Announce Type: new Abstract: Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and proofs. Most current applications
17. Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26203
摘要:arXiv:2606.26203v1 Announce Type: new Abstract: As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparat
18. Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26205
摘要:arXiv:2606.26205v1 Announce Type: new Abstract: Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records, which are authoritat
19. Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26267
摘要:arXiv:2606.26267v1 Announce Type: new Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess. However, they inherently suffer from response lag due to their exclusive relianc
20. Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems
- 来源:arXiv cs.AI
- 发布时间:2026-06-26 04:00 UTC
- 链接:https://arxiv.org/abs/2606.26298
摘要:arXiv:2606.26298v1 Announce Type: new Abstract: Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment. This paper observes that huma