发布日期:2026-09-08
收录条目:20
1. OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device
- 来源:MarkTechPost
- 发布时间:2026-09-07 19:18 UTC
- 链接:https://www.marktechpost.com/2026/09/07/openbmb-releases-minicpm5-2b-a-2-52b-dense-model-averaging-53-9-across-34-benchmarks-and-built-to-run-on-device/
摘要:OpenBMB has released MiniCPM5-2B, a dense causal language model with 2,516,756,480 parameters and a native 131,072 token context. It averages 53.9 across the 34 benchmarks in its model card, ahead of Qwen3.5-4B at 51.1,
2. Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories
- 来源:MarkTechPost
- 发布时间:2026-09-07 18:38 UTC
- 链接:https://www.marktechpost.com/2026/09/07/axis-robotics-releases-axis-a-browser-based-data-engine-with-207-robot-manipulation-tasks-and-50129-trajectories/
摘要:Robot datasets have grown far slower than the models trained on them, mostly because collection stays locked to lab hardware. AXIS moves demonstration collection into a web browser and pushes everything expensive to back
3. IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B
- 来源:MarkTechPost
- 发布时间:2026-09-07 05:00 UTC
- 链接:https://www.marktechpost.com/2026/09/06/ifm-releases-k2-horizon-six-apache-2-0-models-from-0-9b-to-375b/
摘要:Most open model launches release one checkpoint and a benchmark table. The Institute of Foundation Models (IFM) released something wider last week. IFM is the frontier lab launched by MBZUAI in May 2025. K2 Horizon is a
4. EXAONE Forecast for Finance
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04239
摘要:arXiv:2609.04239v1 Announce Type: new Abstract: This technical report presents EXAONE Forecast for Finance (EXAONE Finance), a financial time series (TS) foundation model (TSFM) tailored to financial forecasting. Recent
5. From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04286
摘要:arXiv:2609.04286v1 Announce Type: new Abstract: Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare c
6. Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04298
摘要:arXiv:2609.04298v1 Announce Type: new Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Ada
7. Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04300
摘要:arXiv:2609.04300v1 Announce Type: new Abstract: Ensuring the security of the power system is essential for stability and reliability, especially in the event of disruption. Effective classification of contingency in powe
8. Iris: Climbing to the Search Frontier
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04304
摘要:arXiv:2609.04304v1 Announce Type: new Abstract: We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are
9. A Removal Based Approach to Improve LLM Faithfulness at Test-Time
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04343
摘要:arXiv:2609.04343v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, thes
10. Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04373
摘要:arXiv:2609.04373v1 Announce Type: new Abstract: Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving
11. Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04377
摘要:arXiv:2609.04377v1 Announce Type: new Abstract: Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrate encoding how they decide, negotiate, and execute. Generic LL
12. HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04444
摘要:arXiv:2609.04444v1 Announce Type: new Abstract: Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put a price on avoiding the side effect and to name tha
13. PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04476
摘要:arXiv:2609.04476v1 Announce Type: new Abstract: Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, sto
14. When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04490
摘要:arXiv:2609.04490v1 Announce Type: new Abstract: Quantization is widely used to reduce the computational and memory demands of neural-network inference. In recurrent networks, however, the quantized state is stored and re
15. ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04493
摘要:arXiv:2609.04493v1 Announce Type: new Abstract: We present ResLearn-XR, a residual learning framework for predicting eXtended Reality (XR) network traffic and estimating Quality-of-Experience (QoE) risk. ResLearn-XR adop
16. Rethinking Indirect Prompt Injection as a Test-Time Search Problem
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04495
摘要:arXiv:2609.04495v1 Announce Type: new Abstract: We formulate indirect prompt injection as a test-time search over a task-dependent attack surface induced by the environment, user task, and injection task. To operationali
17. BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04504
摘要:arXiv:2609.04504v1 Announce Type: new Abstract: Cardiac, neural, behavioral, and speech measurements from wearable and mobile devices provide partial, noise-sensitive views of physiological state. BioSync combines these
18. What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04518
摘要:arXiv:2609.04518v1 Announce Type: new Abstract: Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness recipe mixes two choices: exposing the policy to several harnesses
19. MaxKernel: Agentic Kernel Generation for TPUs
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04523
摘要:arXiv:2609.04523v1 Announce Type: new Abstract: Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be l
20. Towards a universal language of concepts: A survey
- 来源:arXiv cs.AI
- 发布时间:2026-09-07 04:00 UTC
- 链接:https://arxiv.org/abs/2609.04528
摘要:arXiv:2609.04528v1 Announce Type: new Abstract: Humans can learn and generalize novel concepts from sparse data because they express knowledge in rich structural formats. In this paper, we propose that programs are a str