| 1 |
SWE-Bench Pro Public Dataset / Scale AI |
A |
沿用既有阅读记录 |
| 2 |
SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution |
A |
沿用既有阅读记录 |
| 3 |
LongCat-2.0 · 来源 2 |
A |
沿用既有阅读记录 |
| 4 |
AReaL 2.0 / Next-Generation Agentic RL Systems · 来源 2 |
A |
沿用既有阅读记录 |
| 5 |
AutoMem: Automated Learning of Memory as a Cognitive Skill |
A |
沿用既有阅读记录 |
| 6 |
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity |
A |
沿用既有阅读记录 |
| 7 |
NVIDIA ASPIRE: Agentic Skills Discovery for Robotics |
A |
沿用既有阅读记录 |
| 8 |
Safari MCP server for web developers |
A |
沿用既有阅读记录 |
| 9 |
Manufact MCP Cloud / mcp-use production lifecycle |
A |
沿用既有阅读记录 |
| 10 |
Cloudflare AI traffic options: Search / Agent / Training crawler split |
A |
沿用既有阅读记录 |
| 11 |
OpenSquilla 0.4.0 coding mode and verification workflow · 来源 2 |
A |
沿用既有阅读记录 |
| 12 |
Raft Blog:Agent-native workspace / AX / human-agent team product language |
A |
沿用既有阅读记录 |
| 13 |
AutoPass:Evidence-Guided LLM Agents for Compiler Performance Tuning |
A |
沿用既有阅读记录 |
| 14 |
PowerAgentBench-Dyn:A Benchmark for Agentic AI in Power System Dynamic Studies |
A |
沿用既有阅读记录 |
| 15 |
RetailBench:Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments |
A |
沿用既有阅读记录 |
| 16 |
Multi-LCB:Extending LiveCodeBench to Multiple Programming Languages |
A |
沿用既有阅读记录 |
| 17 |
Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference |
A |
沿用既有阅读记录 |
| 18 |
Apodex-1.0:verification-centric deep-research agent team / AgentOS / AgentHarness · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 19 |
zartbot:用 Agentic Workflow 做大模型全栈研究 / AI-infra auto-research |
A |
沿用既有阅读记录 |
| 20 |
Oriol Vinyals / Gemini Co-Lead:World Models、Agent Scaffolding、Memory 与 Post-Training RL 路线访谈 |
A |
沿用既有阅读记录 |
| 21 |
MemDecoder:把 memory composition 做成 autoregressive index decoding |
A |
沿用既有阅读记录 |
| 22 |
MemEvolve:把 agent memory architecture 本身纳入 meta-evolution · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 23 |
SimpleMem:semantic lossless compression + adaptive retrieval 的 lifelong memory substrate · 来源 2 · 来源 3 · 来源 4 |
A |
沿用既有阅读记录 |
| 24 |
MemOCR:把 memory 渲染成 layout-aware visual context 来分配信息密度 · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 25 |
Darwinian Memory:GUI agent 的 utility-driven natural selection memory · 来源 2 |
A |
沿用既有阅读记录 |
| 26 |
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective · 来源 2 |
A |
沿用既有阅读记录 |
| 27 |
MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems |
A |
沿用既有阅读记录 |
| 28 |
Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents |
A |
沿用既有阅读记录 |
| 29 |
SemiAnalysis / 腾讯科技:AI Dark Output 与 GDP 统计黑洞 |
A |
沿用既有阅读记录 |
| 30 |
ACE: Agentic Context Engineering for Self-Improving Language Models |
A |
沿用既有阅读记录 |
| 31 |
ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory |
A |
沿用既有阅读记录 |
| 32 |
XENON: Experience-based Knowledge Correction for Robust Planning in Minecraft |
A |
沿用既有阅读记录 |
| 33 |
FlowSearcher: Synthesizing Memory-Guided Agentic Workflows for Web Information Seeking |
A |
沿用既有阅读记录 |
| 34 |
AgentFlow: In-the-Flow Agentic System Optimization for Effective Planning and Tool Use |
A |
沿用既有阅读记录 |
| 35 |
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent |
A |
沿用既有阅读记录 |
| 36 |
REMem: Reasoning with Episodic Memory in Language Agent |
A |
沿用既有阅读记录 |
| 37 |
ReMemR1: Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents |
A |
沿用既有阅读记录 |
| 38 |
GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs |
A |
沿用既有阅读记录 |
| 39 |
GLoW: Dual-Scale World Memory for LLM Agents towards Hard-Exploration Problems |
A |
沿用既有阅读记录 |
| 40 |
Improving Code Localization with Repository Memory |
A |
沿用既有阅读记录 |
| 41 |
TokMem: One-Token Procedural Memory for Large Language Models |
A |
沿用既有阅读记录 |
| 42 |
Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management |
A |
沿用既有阅读记录 |
| 43 |
DataMind: Scaling Generalist Data-Analytic Agents |
A |
沿用既有阅读记录 |
| 44 |
MemGen: Weaving Generative Latent Memory for Self-Evolving Agents |
A |
沿用既有阅读记录 |
| 45 |
Agent-World:可扩展真实环境合成与自进化 Agent 训练 |
A |
公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 |
| 46 |
LangMem / LangGraph long-term memory docs |
B |
沿用既有阅读记录 |
| 47 |
Microsoft GraphRAG / graph-based retrieval baseline |
B |
沿用既有阅读记录 |
| 48 |
Hands-On Modern RL |
A |
沿用既有阅读记录 |
| 49 |
AgentLongBench:用 environment rollout 测 long-context agent,而不是静态 retrieval · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 50 |
PlugMem:把 raw experience 压成 knowledge-centric memory graph · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 51 |
SGLang Disaggregated Serving:P/D split 与 KV transfer 的另一套工程参照 |
B |
沿用既有阅读记录 |
| 52 |
Continuum:multi-turn agent scheduling with KV cache TTL · 来源 2 |
A |
沿用既有阅读记录 |
| 53 |
Agentic KV-cache radar:PrefillShare / TokenDance / CONCUR · 来源 2 · 来源 3 |
B |
沿用既有阅读记录 |
| 54 |
ThunderAgent:program-aware agentic inference system · 来源 2 |
A |
沿用既有阅读记录 |
| 55 |
Cloudflare Code Mode MCP:把工具面压成可执行代码接口 · 来源 2 |
A |
沿用既有阅读记录 |
| 56 |
OpenAI Responses API WebSocket:agent loop state reuse as API/runtime optimization |
A |
沿用既有阅读记录 |
| 57 |
Don't Break the Cache:long-horizon agent prompt caching 的策略边界 |
A |
沿用既有阅读记录 |
| 58 |
Datadog State of AI Engineering:生产 agent 的 prompt/cache/observability 统计锚点 |
A |
沿用既有阅读记录 |
| 59 |
NVIDIA Extreme Co-Design for Agentic Systems:agent token economics / sub-agent / compaction 公开案例 |
A |
沿用既有阅读记录 |
| 60 |
LangSmith / Langfuse observability model:trace / run / session / feedback / observation |
A |
沿用既有阅读记录 |
| 61 |
AgentTrace / AgentSight:agent observability beyond LLM tracing · 来源 2 · 来源 3 |
B |
沿用既有阅读记录 |
| 62 |
Lilian Weng: Why We Think |
A |
沿用既有阅读记录 |
| 63 |
ezyang: OSS code review, in the era of LLMs |
A |
沿用既有阅读记录 |
| 64 |
Workflow-aware serving radar:Helium / ForkKV / ToolCacheAgent · 来源 2 · 来源 3 |
B |
沿用既有阅读记录 |
| 65 |
GPU MODE Resource Stream / Reference Kernels · 来源 2 · 来源 3 |
B |
沿用既有阅读记录 |
| 66 |
PyTorch Dev Discuss: performance / deployment categories |
B |
沿用既有阅读记录 |
| 67 |
Claw-Eval-Live:用 live workflow demand signal 校准 agent benchmark · 来源 2 |
A |
沿用既有阅读记录 |
| 68 |
Claw-Eval:Pass@k / Pass^k 与 trajectory-aware grading 的 agent eval 锚点 · 来源 2 |
A |
沿用既有阅读记录 |
| 69 |
AgentSwing:parallel branch + lookahead routing 的 context management 机制线索 |
A |
沿用既有阅读记录 |
| 70 |
SkillFlow:lifelong skill discovery / repair / library evolution |
A |
沿用既有阅读记录 |
| 71 |
Live-Evo:online self-evolving memory from continuous feedback |
A |
沿用既有阅读记录 |
| 72 |
MemSkill:learnable and evolvable memory skills |
A |
沿用既有阅读记录 |
| 73 |
MemoryCD:cross-domain user memory benchmark from real behavior |
A |
沿用既有阅读记录 |
| 74 |
A-Evolve:agent evolution as infrastructure · 来源 2 |
B |
沿用既有阅读记录 |
| 75 |
agent-eval:Inspect-format eval suite and leaderboard tooling |
B |
沿用既有阅读记录 |
| 76 |
Anthropic: Harness design for long-running application development |
S |
沿用既有阅读记录 |
| 77 |
10 篇论文拆解 Skill + 自进化的技术路线 |
S |
沿用既有阅读记录 |
| 78 |
火山养“龙虾”日志:Self-Improving Skill 让 AI 学会“自我进化” |
S |
沿用既有阅读记录 |
| 79 |
Context Engineering 2.0: The Context of Context Engineering · 来源 2 |
A |
沿用既有阅读记录 |
| 80 |
Anthropic: Effective context engineering for AI agents |
A |
沿用既有阅读记录 |
| 81 |
Prompting Guide: Context Engineering Guide |
B |
沿用既有阅读记录 |
| 82 |
微信导读:超越 Prompt 和 RAG,「上下文工程」成了 Agent 核心胜负手 |
B |
沿用既有阅读记录 |
| 83 |
Memex(RL):stable index + full-fidelity evidence 的 indexed experience memory |
S |
沿用既有阅读记录 |
| 84 |
AgentRR:Record & Replay 范式、multi-level experience 与 check function |
A |
沿用既有阅读记录 |
| 85 |
A-MEM:Agentic Memory 的动态索引、链接与 memory evolution · 来源 2 |
S |
沿用既有阅读记录 |
| 86 |
Mem0:生产级 agent memory 的 hybrid search、entity linking 与 token/latency 约束 · 来源 2 |
A |
沿用既有阅读记录 |
| 87 |
ProRAG:RAG 中 process reward 与 step-level credit assignment |
A |
沿用既有阅读记录 |
| 88 |
AgeMem:把长期/短期记忆管理整合进 agent policy |
A |
沿用既有阅读记录 |
| 89 |
Mem-T:Memory Operation Tree 与 MoT-GRPO 的 reward densification |
A |
沿用既有阅读记录 |
| 90 |
UMA / Learning to Remember:end-to-end memory agent、CRUD 与 Ledger-QA |
A |
沿用既有阅读记录 |
| 91 |
Memory-R1:用 RL 学习 ADD/UPDATE/DELETE/NOOP 与 memory utilization |
A |
沿用既有阅读记录 |
| 92 |
OCR-Memory:用视觉锚点保存长程 agent 轨迹并按需转录原文 |
A |
沿用既有阅读记录 |
| 93 |
MemOS:把 memory 当作可管理系统资源的 Memory OS · 来源 2 |
A |
沿用既有阅读记录 |
| 94 |
Engram:DeepSeek 条件记忆,把静态查表从动态推理中拆出来 · 来源 2 |
B |
沿用既有阅读记录 |
| 95 |
LLMs-augmented Contextual Bandit:LLM encoder + bandit 的基础概念锚点 |
B |
沿用既有阅读记录 |
| 96 |
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks · 来源 2 |
S |
沿用既有阅读记录 |
| 97 |
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling |
S |
沿用既有阅读记录 |
| 98 |
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents |
S |
沿用既有阅读记录 |
| 99 |
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees |
S |
沿用既有阅读记录 |
| 100 |
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices |
S |
沿用既有阅读记录 |
| 101 |
Welcome to the Era of Experience:Sutton / Silver 的经验时代与 Agentic RL 主张 |
A |
沿用既有阅读记录 |
| 102 |
DeepResearcher:真实 Web 环境中的 deep research agent RL · 来源 2 |
A |
公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 |
| 103 |
Rethinking Memory in LLM-based Agents:AI 记忆表示、六大原子操作与 Memory Compass · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 104 |
WebXSkill + Browserbase Skills:可执行 Skill 作为 Web Agent substrate · 来源 2 |
S |
沿用既有阅读记录 |
| 105 |
τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment |
A |
沿用既有阅读记录 |
| 106 |
agentevals:基于 OpenTelemetry trace 的本地离线 Agent 评测 |
A |
沿用既有阅读记录 |
| 107 |
BudgetMem: Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory |
A |
沿用既有阅读记录 |
| 108 |
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents |
A |
沿用既有阅读记录 |
| 109 |
ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving |
A |
沿用既有阅读记录 |
| 110 |
AgentTrace: A Structured Logging Framework for Agent System Observability |
A |
沿用既有阅读记录 |
| 111 |
o11y-bench:AI Agent 的 Observability 任务 benchmark |
A |
沿用既有阅读记录 |
| 112 |
ATBench:trajectory-level agent safety benchmark |
A |
沿用既有阅读记录 |
| 113 |
mem9:OpenClaw 云端永续记忆、Memory Space 与 ContextEngine 生命周期接口 · 来源 2 |
A |
沿用既有阅读记录 |
| 114 |
ArkClaw 十大热门 Skills:技能生态、加载优先级、多 Agent 绑定与 APM 观测 |
A |
沿用既有阅读记录 |
| 115 |
How To Be A World-Class Agentic Engineer:保持简单、上下文管理与 contract-driven session |
A |
沿用既有阅读记录 |
| 116 |
Agent Skills Prompt Injection:Skill 文件带来的现实攻击面 |
A |
沿用既有阅读记录 |
| 117 |
Claw Code / Claude Code 开源 Rust agent harness 解读 |
A |
沿用既有阅读记录 |
| 118 |
Multi-Agent 是伪技术路线?OpenClaw subagents 与多 Agent 适用边界 · 来源 2 |
A |
沿用既有阅读记录 |
| 119 |
Why Do Multi-Agent LLM Systems Fail?:MAST 多智能体失败分类与 trace 数据集 · 来源 2 |
A |
沿用既有阅读记录 |
| 120 |
Claude Code Superpowers:结构化软件工程技能框架 · 来源 2 |
A |
沿用既有阅读记录 |
| 121 |
EvoMap / Evolver:GEP 能力进化资产与跨 Agent 经验共享 |
A |
沿用既有阅读记录 |
| 122 |
Agent Harness Engineering: A Survey · 来源 2 · 来源 3 |
A |
沿用既有阅读记录;旧记录含已读备注但受管生命周期仍为 candidate;保留状态冲突,待单独核对。 |
| 123 |
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents |
A |
沿用既有阅读记录 |
| 124 |
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization · 来源 2 |
A |
沿用既有阅读记录 |
| 125 |
SkillOS: Learning Skill Curation for Self-Evolving Agents · 来源 2 |
A |
沿用既有阅读记录 |
| 126 |
Voyager: An Open-Ended Embodied Agent with Large Language Models · 来源 2 |
A |
沿用既有阅读记录 |
| 127 |
Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers |
A |
沿用既有阅读记录 |
| 128 |
AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems |
A |
沿用既有阅读记录 |
| 129 |
Coding Agents are Effective Long-Context Processors |
A |
沿用既有阅读记录 |
| 130 |
On Data Engineering for Scaling LLM Terminal Capabilities:Terminal Agent 数据工程 · 来源 2 |
A |
沿用既有阅读记录 |
| 131 |
A Guide to Gen AI / LLM Vibecoding for Expert Programmers |
A |
沿用既有阅读记录 |
| 132 |
Exgentic Open Agent Leaderboard:跨 benchmark 的统一 Agent 评测协议 |
A |
沿用既有阅读记录 |
| 133 |
AgencyBench:1M-token long-horizon autonomous agent benchmark |
A |
沿用既有阅读记录 |
| 134 |
Agent Replay:local-first desktop evals / observability / memory |
B |
沿用既有阅读记录 |
| 135 |
OpenAI: Speeding up agentic workflows with WebSockets in the Responses API |
A |
沿用既有阅读记录 |
| 136 |
Glean Waldo:专门的 agentic search model 做检索规划 |
A |
沿用既有阅读记录 |
| 137 |
CC Pocket |
A |
沿用既有阅读记录 |
| 138 |
ArkClaw 动手实验上新:财经、评论洞察、财报、视频剪辑、播客生成 |
A |
沿用既有阅读记录 |
| 139 |
火山联网搜索 Skill:ArkClaw / OpenClaw 的 Agent 原生搜索能力 |
A |
沿用既有阅读记录 |
| 140 |
被 OpenClaw 推上风口的飞书:企业 Agent 落地、工作空间容器与组织协作变迁 |
A |
沿用既有阅读记录 |
| 141 |
碎片化编程与手机访问 Claude Code / OpenClaw |
A |
沿用既有阅读记录 |
| 142 |
Agent-Reach:开源本地互联网工具脚手架 |
A |
沿用既有阅读记录 |
| 143 |
Amazon SynerGen:搜推召排一体化的生成式推荐框架 · 来源 2 |
A |
沿用既有阅读记录 |
| 144 |
快手 TagCF:用户角色、行为逻辑图与推荐系统显式破茧 · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 145 |
DualGR:快手长短期兴趣建模的生成式召回实践 · 来源 2 |
A |
沿用既有阅读记录 |
| 146 |
LEMUR:大规模端到端多模态推荐 |
A |
沿用既有阅读记录 |
| 147 |
Don’t Waste It:用结构化 Human Priors 指导生成式推荐多头解码 · 来源 2 |
A |
沿用既有阅读记录 |
| 148 |
SSCTL:多领域推荐中的数据不均衡与半监督迁移 |
A |
沿用既有阅读记录 |
| 149 |
GNOLR:多隐式反馈的有序偏好统一嵌入与推荐召回简化 · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 150 |
搜索推荐推理合一 / LLM 混合 ID 生成:NEO at Spotify · 来源 2 |
A |
沿用既有阅读记录 |
| 151 |
Google STATIC:向量化 Trie,加速生成式召回约束解码 · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 152 |
CollectiveKV:序列推荐中的跨用户 KV cache 共享与压缩 · 来源 2 |
A |
沿用既有阅读记录 |
| 153 |
Hiformer:Google Play 推荐排序中的异构特征交互 Transformer · 来源 2 |
A |
沿用既有阅读记录 |
| 154 |
HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction |
A |
沿用既有阅读记录 |
| 155 |
X / Twitter For You 推荐算法:Grok-based Transformer、两塔召回与 Candidate Isolation · 来源 2 |
A |
沿用既有阅读记录 |
| 156 |
Meta SilverTorch:召回与粗排全链路 GPU 算子化 · 来源 2 |
A |
沿用既有阅读记录 |
| 157 |
LORE:阿里广告搜索相关性大模型与垂直 RL 经验 · 来源 2 |
A |
沿用既有阅读记录 |
| 158 |
谈谈 RL Infra / FlashRL:训练、推理、调度、权重同步与 observability · 来源 2 |
A |
沿用既有阅读记录 |
| 159 |
Ilya Sutskever:An Observation on Generalization / 用压缩视角理解无监督学习 |
A |
沿用既有阅读记录 |
| 160 |
Muon 优化器:矩阵正交化更新、LLM 训练可扩展性与 Kimi K2 的 MuonClip · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 161 |
DeepSeek-V4 Technical Report:Million-Token Context、CSA/HCA、mHC、Muon 与长程 Agent 能力 · 来源 2 |
A |
沿用既有阅读记录 |
| 162 |
PyGraph:让 PyTorch 的 CUDA Graph 优化更高效 · 来源 2 |
A |
沿用既有阅读记录 |
| 163 |
PyTorch FlexAttention:用 score_mod / mask_mod 表达 Attention 变体并编译成高性能 Triton kernel · 来源 2 |
A |
沿用既有阅读记录 |
| 164 |
Intel / 龙蜥:xFasterTransformer、至强 AMX/HBM/CXL 与 CPU LLM 推理背景 |
A |
沿用既有阅读记录 |
| 165 |
Look Ma, No Bubbles!:Llama-1B 低延迟 Megakernel 与 memory pipeline bubbles · 来源 2 |
A |
沿用既有阅读记录 |
| 166 |
CUDA Agent:基于真实性能反馈的 agentic RL CUDA 内核生成 · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 167 |
AutoResearch 自研高性能 GPU 算子 Flash Attention:MFU 42%、Opus 4.6、8 小时 25 轮迭代 |
A |
沿用既有阅读记录 |
| 168 |
DualPath:Agentic LLM Inference 的 KV-Cache 存储带宽瓶颈 |
A |
沿用既有阅读记录 |
| 169 |
MetaShuffling:Llama 4 MoE 推理的 shuffle-based Fused MoE kernel 与 Padding 避免 |
A |
沿用既有阅读记录 |
| 170 |
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training? |
A |
沿用既有阅读记录 |
| 171 |
少用 sense 挑战 math:如何把 post train 做好 |
A |
沿用既有阅读记录 |
| 172 |
TTT-E2E:End-to-End Test-Time Training for Long Context · 来源 2 |
A |
沿用既有阅读记录 |
| 173 |
Nested Learning / Hope:多频率持续学习、优化器即联想记忆 · 来源 2 |
A |
沿用既有阅读记录 |
| 174 |
林俊旸详解 Qwen 模型设计中的每一次取舍 · 来源 2 |
A |
沿用既有阅读记录 |
| 175 |
mHC: Manifold-Constrained Hyper-Connections |
A |
公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 |
| 176 |
MiniMax VTP:Towards Scalable Pre-training of Visual Tokenizers for Generation · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 177 |
Gated Attention for Large Language Models:Non-linearity, Sparsity, and Attention-Sink-Free |
A |
公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 |
| 178 |
Over-Tokenized Transformer:扩大输入词表作为稀疏 scaling 维度 |
A |
沿用既有阅读记录 |
| 179 |
MiniMax M2:为什么最终选择 full attention |
A |
沿用既有阅读记录 |
| 180 |
王云鹤:Harness 是复杂优化问题,Agent = Models + Harness |
A |
沿用既有阅读记录 |
| 181 |
田渊栋访谈:Meta 裁员后,对 AI 研究、RL 与职业选择的思考 |
A |
沿用既有阅读记录 |
| 182 |
WhynotTV Podcast #4:翁家翌谈 OpenAI、post-training、RL infra、工程能力与 impact |
S |
沿用既有阅读记录 |
| 183 |
田渊栋:告别 OpenAI 小作文(二)未来会是什么样子 |
S |
沿用既有阅读记录 |
| 184 |
Manus季逸超访谈解读(1):Benchmark是所有AI公司的唯一护城河 |
S |
沿用既有阅读记录 |
| 185 |
小红书:从 L3 到 L8 分别需要怎么 vibe |
A |
沿用既有阅读记录 |
| 186 |
小红书:张咋啦 follow-builders skill / 关注 builders 而不是 KOL |
A |
沿用既有阅读记录 |
| 187 |
Aha:AI 产品真正杠杆在后端,而不是对话框 |
A |
沿用既有阅读记录 |
| 188 |
如何做出好产品:张一鸣产品定律、推荐、AB 测试与务实文化 |
A |
沿用既有阅读记录 |
| 189 |
对话腾讯 ima 产品团队:有价值的产品,不需要告诉用户「这是智能体」 |
A |
沿用既有阅读记录 |
| 190 |
锦秋基金臧天宇:我们看 AI 应用的思考脉络 |
A |
沿用既有阅读记录 |
| 191 |
Agent 经济学:一人公司、交易成本、Evals 与管理成本转移 |
B |
沿用既有阅读记录 |
| 192 |
Anthropic CEO Dario Amodei 访谈:Scaling Laws、动态数据、应用层机会与 AI 时代职业判断 |
B |
沿用既有阅读记录 |
| 193 |
DeepMind AlphaGo 十年复盘:游戏化、搜索、验证器与科学发现 |
B |
沿用既有阅读记录 |
| 194 |
Werner Vogels final keynote:Renaissance Developer 与 AI 时代工程师价值 |
B |
沿用既有阅读记录 |
| 195 |
SaaS + Agent 十人谈:盖雅工场的垂直 Agent 与结果付费判断 |
B |
沿用既有阅读记录 |
| 196 |
AI Agent 产品落地五个坑:记忆、工具幻觉、调度、成本与执行精度 |
B |
沿用既有阅读记录 |
| 197 |
Stripe Sessions 2026:Agentic commerce 与智能体支付基础设施 |
B |
沿用既有阅读记录 |
| 198 |
胡渊鸣:生成式 AI 这门生意(上) |
B |
沿用既有阅读记录 |
| 199 |
Nicholas Wilt 推荐的计算机书单 |
B |
沿用既有阅读记录 |
| 200 |
小红书:为什么人类活不过二百岁 |
B |
沿用既有阅读记录 |
| 201 |
大模型第一性原理:信息论、定向信息与 token 信道抽象 |
B |
沿用既有阅读记录 |
| 202 |
Artem Kirsanov:概率背后的核心方程,熵、交叉熵与 KL 散度 |
B |
沿用既有阅读记录 |
| 203 |
DataFunTalk:字节和腾讯 Agentic RL 最佳实践议题线索 |
B |
沿用既有阅读记录 |
| 204 |
大模型的第一性原理(一):统计物理篇 |
B |
沿用既有阅读记录 |
| 205 |
Visual Language Hypothesis:为什么自监督难以直接学到语义 · 来源 2 |
B |
沿用既有阅读记录 |
| 206 |
知乎专栏文章:标题暂不可见 |
U |
未读 |
| 207 |
Textual SGD: Optimizing External Memory |
U |
未读 |
| 208 |
Static AI to Continual Enterprise Learning: A Living Dialect of Tribal Knowledge |
S |
沿用既有阅读记录 |
| 209 |
Sema Code: Decoupling AI Coding Agents from the Terminal |
S |
沿用既有阅读记录 |
| 210 |
Thinking with Visual Primitives |
A |
沿用既有阅读记录 |
| 211 |
Microsoft Agent 365 GA |
A |
沿用既有阅读记录 |
| 212 |
Introducing Cloudflare Agent Cloud |
A |
沿用既有阅读记录 |
| 213 |
Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructure |
S |
沿用既有阅读记录 |
| 214 |
MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents |
S |
沿用既有阅读记录 |
| 215 |
Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers |
A |
沿用既有阅读记录 |
| 216 |
Joel Leibo 研究脉络:multi-agent 的对象不是单个 agent,而是关系结构 |
A |
沿用既有阅读记录 |
| 217 |
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces |
S |
沿用既有阅读记录 |
| 218 |
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills |
S |
沿用既有阅读记录 |
| 219 |
GPT-5.5 Instant: smarter, clearer, and more personalized |
A |
沿用既有阅读记录 |
| 220 |
Gemini API File Search is now multimodal |
A |
沿用既有阅读记录 |
| 221 |
How OpenAI delivers low-latency voice AI at scale |
A |
沿用既有阅读记录 |
| 222 |
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection |
A |
沿用既有阅读记录 |
| 223 |
OpenAI MRC: Supercomputer networking to accelerate large scale AI training |
S |
沿用既有阅读记录 |
| 224 |
Agents can now create Cloudflare accounts, buy domains, and deploy |
S |
沿用既有阅读记录 |
| 225 |
Anthropic: Higher usage limits for Claude and a compute deal with SpaceX |
A |
沿用既有阅读记录 |
| 226 |
Gemini API File Search is now multimodal |
A |
沿用既有阅读记录 |
| 227 |
Gemini API Webhooks for long-running jobs |
A |
沿用既有阅读记录 |
| 228 |
CopilotKit raises $27M Series A |
A |
沿用既有阅读记录 |
| 229 |
Model-Harness-Fit:同一模型换壳性能会显著漂移 |
A |
沿用既有阅读记录 |
| 230 |
Anthropic: Natural Language Autoencoders: Turning Claude's Thoughts into Text · 来源 2 |
S |
沿用既有阅读记录 |
| 231 |
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces · 来源 2 |
S |
沿用既有阅读记录 |
| 232 |
Do Agent Rules Shape or Distort? Guardrails Beat Guidance in Coding Agents |
A |
沿用既有阅读记录 |
| 233 |
agent-skills-eval |
A |
沿用既有阅读记录 |
| 234 |
Accelerating Gemma 4: faster inference with multi-token prediction drafters |
A |
沿用既有阅读记录 |
| 235 |
ds4: DeepSeek 4 Flash local inference engine for Metal |
A |
沿用既有阅读记录 |
| 236 |
OpenAI: Advancing voice intelligence with new models in the API |
A |
沿用既有阅读记录 |
| 237 |
AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields |
A |
沿用既有阅读记录 |
| 238 |
NVIDIA Dynamo Agent Context and Tracing |
A |
沿用既有阅读记录 |
| 239 |
Ray A2A multi-agent + Anyscale scalable MCP on Ray Serve |
A |
沿用既有阅读记录 |
| 240 |
OpenAI Codex browser / Chrome workflow signal |
A |
沿用既有阅读记录 |
| 241 |
Claude Microsoft 365 connector and enterprise context permissions |
A |
沿用既有阅读记录 |
| 242 |
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? |
S |
沿用既有阅读记录 |
| 243 |
LLMs Corrupt Your Documents When You Delegate |
S |
沿用既有阅读记录 |
| 244 |
Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents |
S |
沿用既有阅读记录 |
| 245 |
SysMoBench / Specula: Can LLMs model real-world systems in TLA+? · 来源 2 |
A |
沿用既有阅读记录 |
| 246 |
re_gent: Version-Control for AI coding agents |
A |
沿用既有阅读记录 |
| 247 |
A recent experience with ChatGPT 5.5 Pro |
A |
沿用既有阅读记录 |
| 248 |
MiniMax “马嘉祺”问题:后训练 token 覆盖不足导致输出映射层退化 |
A |
沿用既有阅读记录 |
| 249 |
How AI-pilled are you? / P9 AI Fluency Index |
A |
沿用既有阅读记录 |
| 250 |
Jiayi Weng:Learning Beyond Gradients / Heuristic Learning |
S |
沿用既有阅读记录 |
| 251 |
MiniMax:一个 AI 还是不够 / Agent Teams 与 Mavis |
A |
沿用既有阅读记录 |
| 252 |
甲子光年:Kimi 总裁张予彤北大实录:我们想要有抽象能力和偏执的人 |
A |
沿用既有阅读记录 |
| 253 |
MemReranker:Reasoning-Aware Reranking for Agent Memory Retrieval |
S |
沿用既有阅读记录 |
| 254 |
Thinking Machines:Interaction Models |
S |
沿用既有阅读记录 |
| 255 |
OpenAI:Running Codex safely at OpenAI |
S |
沿用既有阅读记录 |
| 256 |
Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning |
S |
沿用既有阅读记录 |
| 257 |
WildClawBench:Real-World Long-Horizon Agent Evaluation |
A |
沿用既有阅读记录 |
| 258 |
Needle:26M function calling model |
A |
沿用既有阅读记录 |
| 259 |
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks · 来源 2 |
S |
沿用既有阅读记录 |
| 260 |
OpenAI Daybreak / Codex Security control plane |
S |
沿用既有阅读记录 |
| 261 |
OpenAI: Investigating the consequences of accidentally grading CoT during RL |
A |
沿用既有阅读记录 |
| 262 |
UiPath for Coding Agents |
A |
沿用既有阅读记录 |
| 263 |
SGLang v0.5.11 release |
A |
沿用既有阅读记录 |
| 264 |
ELF: Embedded Language Flows · 来源 2 |
A |
沿用既有阅读记录 |
| 265 |
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues |
S |
沿用既有阅读记录 |
| 266 |
TencentDB Agent Memory · 来源 2 |
S |
沿用既有阅读记录 |
| 267 |
Needle: 26M function call model |
A |
沿用既有阅读记录 |
| 268 |
Notion Developer Platform |
A |
沿用既有阅读记录 |
| 269 |
Claude Agent SDK credits and agent subscription economics |
A |
沿用既有阅读记录 |
| 270 |
Alexa for Shopping: personalized agentic commerce |
A |
沿用既有阅读记录 |
| 271 |
Meta Incognito Chat with Meta AI |
B |
沿用既有阅读记录 |
| 272 |
Is Grep All You Need? How Agent Harnesses Reshape Agentic Search |
A |
沿用既有阅读记录 |
| 273 |
Qoder 1.0 / Quest Mode:AI IDE -> autonomous agent workbench |
A |
沿用既有阅读记录 |
| 274 |
ChatGPT Personal Finance |
B |
沿用既有阅读记录 |
| 275 |
Grok Build CLI |
B |
沿用既有阅读记录 |
| 276 |
MeMo: Memory as a Model |
A |
沿用既有阅读记录 |
| 277 |
APWA: A Distributed Architecture for Parallelizable Agentic Workflows |
A |
沿用既有阅读记录 |
| 278 |
Ring-2.6-1T:面向 Agent 执行的开源推理模型 |
A |
沿用既有阅读记录 |
| 279 |
Anthropic: Teaching Claude why |
A |
沿用既有阅读记录 |
| 280 |
AGenUI:三端原生 A2UI Renderer |
B |
沿用既有阅读记录 |
| 281 |
TwiSTAR: Think Fast, Think Slow, Then Act |
A |
沿用既有阅读记录 |
| 282 |
小红书 paper radar:4月近两周自觉不错的 paper 合集第三期 |
A |
沿用既有阅读记录 |
| 283 |
ADAS / Automated Design of Agentic Systems:用搜索自动设计 agent system · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 284 |
DGM / Darwin Gödel Machine:开放式自改进 agent · 来源 2 |
A |
沿用既有阅读记录 |
| 285 |
AFlow:自动生成 agentic workflow · 来源 2 |
A |
沿用既有阅读记录 |
| 286 |
SPO / Self-Supervised Prompt Optimization:无人工标签的 prompt 优化 · 来源 2 |
A |
沿用既有阅读记录 |
| 287 |
GEPA / Reflective Prompt Evolution:反思式 prompt evolution · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 288 |
COMPASS:Enhancing Agent Long-Horizon Reasoning with Evolving Context |
A |
沿用既有阅读记录 |
| 289 |
Training-Free Group Relative Policy Optimization |
A |
沿用既有阅读记录 |
| 290 |
Motive Notes:What Makes 5% of AI Agents Actually Work in Production? |
A |
沿用既有阅读记录 |
| 291 |
AgentEvolver:Towards Efficient Self-Evolving Agent System · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 292 |
To The Crazy Ones | 致超级个体 |
A |
沿用既有阅读记录;既往已读文字正文,视频、图片与白板未逐一展开。讨论完整问题所有权、工具权限、直接用户反馈与组织认可如何支持 AI Builder。 |
| 293 |
Codex / Claude Code memory 模式的收益、风险与 token 成本 · 来源 2 |
A |
沿用既有阅读记录 |
| 294 |
SkyRL:full-stack RL library for LLMs · 来源 2 |
A |
沿用既有阅读记录 |
| 295 |
OpenRLHF:high-performance RLHF / GRPO / PPO framework |
A |
沿用既有阅读记录 |
| 296 |
Selective Rollout:Efficient Reinforcement Learning for Long-Horizon LLM Agents |
A |
沿用既有阅读记录 |
| 297 |
HiPER:State Abstraction and Value-Guided Search for Efficient Long-Horizon Agents |
A |
沿用既有阅读记录 |
| 298 |
AgentFly:Extensible and Scalable Reinforcement Learning for LM Agents |
A |
沿用既有阅读记录 |
| 299 |
Code as Agent Harness · 来源 2 |
S |
沿用既有阅读记录 |
| 300 |
SWE-Chain:Benchmarking Coding Agents on Chained Release-Level Package Upgrades |
S |
沿用既有阅读记录 |
| 301 |
AgentTrust:Runtime Safety Evaluation and Interception for AI Agent Tool Use |
A |
沿用既有阅读记录 |
| 302 |
Google I/O 2026:Gemini 3.5 Flash / Antigravity 2.0 / Managed Agents |
A |
沿用既有阅读记录 |
| 303 |
OpenAI / Databricks:GPT-5.5 on OfficeQA Pro |
A |
沿用既有阅读记录 |
| 304 |
Anthropic acquires Stainless:SDK / CLI / MCP server tooling |
A |
沿用既有阅读记录 |
| 305 |
RRFP:A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability |
A |
沿用既有阅读记录 |
| 306 |
DashAttention:Differentiable and Adaptive Sparse Hierarchical Attention |
A |
沿用既有阅读记录 |
| 307 |
Qwen3.7-Max Agent Frontier |
A |
沿用既有阅读记录 |
| 308 |
OpenAI Guaranteed Capacity:模型 API 的 reserved compute contract |
A |
沿用既有阅读记录 |
| 309 |
TIDE:Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload |
A |
沿用既有阅读记录 |
| 310 |
KoRe:Compact Knowledge Representations for Large Language Models |
A |
沿用既有阅读记录 |
| 311 |
OpenAI model disproves Erdős unit-distance conjecture |
A |
沿用既有阅读记录 |
| 312 |
Multi-Stream LLMs:Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs |
A |
沿用既有阅读记录 |
| 313 |
vLLM x PegaFlow:Production-Grade External KV Cache |
A |
沿用既有阅读记录 |
| 314 |
AWS SageMaker OpenAI-compatible API endpoints |
A |
沿用既有阅读记录 |
| 315 |
Google AI Mode Ads:Gemini-powered ad formats in Search |
A |
沿用既有阅读记录 |
| 316 |
Tencent Hy-MT2:translation model family and IFMTBench |
A |
沿用既有阅读记录 |
| 317 |
Cohere Command A+:可私有部署的 open-source enterprise agent model |
A |
沿用既有阅读记录 |
| 318 |
Domain-Camouflaged Injection Attacks:领域伪装注入绕过 agent guard |
A |
沿用既有阅读记录 |
| 319 |
唐杰:关于 long-horizon tasks 的近期思考 |
A |
沿用既有阅读记录 |
| 320 |
Runtime (YC P26):团队级沙盒化 coding agents |
B |
沿用既有阅读记录 |
| 321 |
Claude Code 源码精读:compact 每次压缩时发生了什么 |
A |
沿用既有阅读记录 |
| 322 |
Neuromancer:Claude Code 上下文管理个人学习笔记 |
A |
沿用既有阅读记录 |
| 323 |
MemGym:长程 agent memory 执行评测环境 |
A |
沿用既有阅读记录 |
| 324 |
Anthropic Project Glasswing / Claude Mythos CVD dashboard |
A |
沿用既有阅读记录 |
| 325 |
CODA:Rewriting Transformer Blocks as GEMM-Epilogue Programs |
A |
沿用既有阅读记录 |
| 326 |
KanBots OSS:本地 Kanban 多 agent 编排 · 来源 2 |
B |
沿用既有阅读记录 |
| 327 |
ReAct:reasoning-action-observation 循环的经典起点 |
A |
沿用既有阅读记录 |
| 328 |
Toolformer:模型自监督学会调用工具 |
A |
沿用既有阅读记录 |
| 329 |
API-Bank:tool-augmented LLM 的 API 评测早期基准 |
A |
沿用既有阅读记录 |
| 330 |
Gorilla:面向海量 API 的 tool retrieval / function calling · 来源 2 |
A |
沿用既有阅读记录 |
| 331 |
ToolLLM / ToolBench:大规模真实 API 的工具学习与评测 · 来源 2 |
A |
沿用既有阅读记录 |
| 332 |
WebArena:真实 Web 环境中的 autonomous agent benchmark |
A |
沿用既有阅读记录 |
| 333 |
VisualWebArena:多模态 Web agent 评测 |
A |
沿用既有阅读记录 |
| 334 |
OSWorld:真实桌面环境中的 computer-use agent benchmark · 来源 2 |
A |
沿用既有阅读记录 |
| 335 |
WorkArena:企业知识工作 Web agent benchmark · 来源 2 |
A |
沿用既有阅读记录 |
| 336 |
SWE-bench:真实 GitHub issue 到 patch 的软件工程评测 |
A |
沿用既有阅读记录 |
| 337 |
OpenHands:通用软件开发 agent 平台 · 来源 2 |
A |
沿用既有阅读记录 |
| 338 |
AutoGen:multi-agent conversation framework · 来源 2 |
A |
沿用既有阅读记录 |
| 339 |
LLMCompiler:并行 function calling / tool execution 编排 |
A |
沿用既有阅读记录 |
| 340 |
AgentLens:agent 行为可视分析与 lucky pass 问题 · 来源 2 |
A |
沿用既有阅读记录 |
| 341 |
Contextual Agent Security:面向不同目的的 agent policy |
A |
沿用既有阅读记录 |
| 342 |
Generative Agents:长期记忆驱动的交互式行为模拟 |
A |
沿用既有阅读记录 |
| 343 |
Lost in the Middle:长上下文位置偏置经典问题 |
A |
沿用既有阅读记录 |
| 344 |
MemGPT:把 LLM memory 管理类比为操作系统 · 来源 2 |
A |
沿用既有阅读记录 |
| 345 |
OpenShell:声明式 policy 驱动的 agent sandbox runtime |
A |
沿用既有阅读记录 |
| 346 |
SWE-ReX:coding agent remote execution / sandbox infrastructure |
A |
沿用既有阅读记录 |
| 347 |
ContextForge:MCP / A2A / REST gateway with governance and observability |
A |
沿用既有阅读记录 |
| 348 |
Agent Governance Toolkit:deterministic policy / identity / sandbox / audit before actions |
A |
沿用既有阅读记录 |
| 349 |
Browser Harness:可编辑 CDP browser harness |
A |
沿用既有阅读记录 |
| 350 |
Symphony:ticket-driven orchestration layer for autonomous implementation runs |
A |
沿用既有阅读记录 |
| 351 |
R2E-Gym:从真实 repo issue 构造 executable coding-agent RL environments · 来源 2 |
A |
沿用既有阅读记录 |
| 352 |
Prime Intellect verifiers:LLM RL environments + evals as reusable verifier library |
A |
沿用既有阅读记录 |
| 353 |
Meta-Harness:把 harness design 本身作为 automated search object |
A |
沿用既有阅读记录 |
| 354 |
Anthropic Context Management:tool result clearing and compaction |
A |
沿用既有阅读记录 |
| 355 |
Context Rot:long context 变长时的性能退化 |
A |
沿用既有阅读记录 |
| 356 |
Anthropic: How we built our multi-agent research system |
A |
沿用既有阅读记录 |
| 357 |
LCGuard:Defending Against Latent Communication in Multi-Agent Systems by System-Level KV Cache Sandboxing |
A |
沿用既有阅读记录 |
| 358 |
Claude Code network sandbox bypass reports |
A |
沿用既有阅读记录 |
| 359 |
Cloudflare Agent Infrastructure Stack |
A |
沿用既有阅读记录 |
| 360 |
Reasonix:prefix-cache-aware terminal coding agent |
A |
沿用既有阅读记录 |
| 361 |
FAME:Fault-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection |
A |
沿用既有阅读记录 |
| 362 |
Epoch AI: AI chip component cost shares |
A |
沿用既有阅读记录 |
| 363 |
Kung & Robinson: On Optimistic Methods for Concurrency Control |
A |
沿用既有阅读记录 |
| 364 |
Shapiro et al.: A comprehensive study of convergent and commutative replicated data types(CRDTs) |
A |
沿用既有阅读记录 |
| 365 |
AeSlides:通过可验证奖励强化幻灯片生成 · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 366 |
ACC: Compiling Agent Trajectories for Long-Context Training |
A |
沿用既有阅读记录 |
| 367 |
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows? · 来源 2 |
A |
沿用既有阅读记录 |
| 368 |
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows · 来源 2 |
A |
沿用既有阅读记录 |
| 369 |
Microsoft Copilot Cowork Exfiltrates Files |
A |
沿用既有阅读记录 |
| 370 |
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation |
A |
沿用既有阅读记录 |
| 371 |
CVEvolve: Autonomous Algorithm Discovery for Unstructured Scientific Data Processing |
A |
沿用既有阅读记录 |
| 372 |
Gemini app becomes more agentic, delivering proactive 24/7 help |
B |
沿用既有阅读记录 |
| 373 |
Anthropic: How we contain Claude across products |
A |
沿用既有阅读记录 |
| 374 |
Language Models Need Sleep |
A |
沿用既有阅读记录 |
| 375 |
EAGLE 3.1: Advancing Speculative Decoding Through Collaboration Between EAGLE, vLLM, and TorchSpec |
A |
沿用既有阅读记录 |
| 376 |
Robin: A multi-agent system for automating scientific discovery |
A |
沿用既有阅读记录 |
| 377 |
Alipay AI Wallet / Token Pay / Agentic Commerce Trust Protocol |
A |
沿用既有阅读记录 |
| 378 |
OpenRouter Raises $113M Series B |
A |
沿用既有阅读记录 |
| 379 |
Xiaomi MiMo-V2.5 Series Price Adjustment |
A |
沿用既有阅读记录 |
| 380 |
Minicor: managed self-healing desktop automation at scale |
A |
沿用既有阅读记录 |
| 381 |
中国企业家:6个月融25亿元,他是“字节系”最猛的AI创业者 |
A |
沿用既有阅读记录 |
| 382 |
ArkClaw 漫剧虾工作流实测:从一个主题到爆款漫剧成片 |
A |
沿用既有阅读记录 |
| 383 |
视频生成 agent / 短剧工具竞品池:Flova / 纳米短剧 / 巨日禄 / 万镜一刻 |
A |
沿用既有阅读记录 |
| 384 |
火山引擎:Vibe Creating,让视频创作回归表达本身 |
A |
沿用既有阅读记录 |
| 385 |
MUSE-Autoskill:Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 386 |
QUEST:Training Frontier Deep Research Agents with Fully Synthetic Tasks |
A |
沿用既有阅读记录 |
| 387 |
Cognition: More Devins in More Places |
A |
沿用既有阅读记录 |
| 388 |
jxnlco: Getting the most out of Codex |
A |
沿用既有阅读记录 |
| 389 |
Polar: Agentic RL on Any Harness at Scale · 来源 2 |
A |
沿用既有阅读记录 |
| 390 |
Structured Agent Distillation for Large Language Model |
A |
沿用既有阅读记录 |
| 391 |
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective |
A |
沿用既有阅读记录 |
| 392 |
Lenz Research: Beyond Benchmarks, Frontier LLM Disagreement on Fact-Checks |
A |
沿用既有阅读记录 |
| 393 |
DBOS: Postgres is All You Need for Durable Workflows |
A |
沿用既有阅读记录 |
| 394 |
OpenAI: How OpenAI uses Codex |
A |
沿用既有阅读记录 |
| 395 |
ClickHouse Agents + Langfuse V4:agentic data stack and observability |
A |
沿用既有阅读记录 |
| 396 |
Claude Code vs Codex scientific-computing head-to-head |
A |
沿用既有阅读记录 |
| 397 |
Coding Beyond Your Training: Claude Code and the Technological Frontier of Software Developers |
A |
沿用既有阅读记录 |
| 398 |
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? |
A |
沿用既有阅读记录 |
| 399 |
LLMSurgeon: Diagnosing Data Mixture of Large Language Models |
A |
沿用既有阅读记录 |
| 400 |
In-Context Reward Adaptation for Robust Preference Modeling |
A |
沿用既有阅读记录 |
| 401 |
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players |
A |
沿用既有阅读记录 |
| 402 |
WALL-WM:World Action Model at Event Boundaries · 来源 2 |
A |
沿用既有阅读记录 |
| 403 |
ATLAS - Autoformalized Textbook Library At Scale |
A |
沿用既有阅读记录 |
| 404 |
tiny-vLLM:Build your own high performance LLM inference engine in C++ and CUDA |
A |
沿用既有阅读记录 |
| 405 |
Never Stop Learning: Continual Learning and Self-Iteration in LLMs |
A |
沿用既有阅读记录 |
| 406 |
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents |
A |
沿用既有阅读记录 |
| 407 |
Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory |
A |
沿用既有阅读记录 |
| 408 |
GodeX:OpenAI Responses API 兼容网关与 provider bridge |
A |
沿用既有阅读记录 |
| 409 |
Meta AI support / Instagram account recovery exploit report |
A |
沿用既有阅读记录 |
| 410 |
Microsoft MAI / Frontier Tuning:workflow-specific RLE 与 Copilot harness model |
A |
沿用既有阅读记录 |
| 411 |
OpenAI Codex for every role / Sites / annotations |
A |
沿用既有阅读记录 |
| 412 |
Microsoft Scout / WorkIQ / Agent 365:always-on work agent control plane |
A |
沿用既有阅读记录 |
| 413 |
Bernini: Latent Semantic Planning for Video Diffusion |
A |
沿用既有阅读记录 |
| 414 |
AdaCodec: A Predictive Visual Code for Video MLLMs |
A |
沿用既有阅读记录 |
| 415 |
CLI-Anything: Towards Agent-Native Computer Use · 来源 2 |
A |
沿用既有阅读记录 |
| 416 |
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management · 来源 2 |
A |
沿用既有阅读记录 |
| 417 |
Taiji: Pareto Optimal Policy Optimization with Semantics-IDs Trade-off for Industrial LLM-Enhanced Recommendation |
A |
沿用既有阅读记录 |
| 418 |
Bringing up DeepSeek-V4-Flash on AMD MI300X · 来源 2 |
B |
沿用既有阅读记录 |
| 419 |
Uber AI coding budget cap / enterprise agent FinOps |
B |
沿用既有阅读记录 |
| 420 |
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents |
A |
沿用既有阅读记录 |
| 421 |
KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks · 来源 2 |
A |
沿用既有阅读记录 |
| 422 |
Multi-Segment Attention / AsymCache:面向 agent serving 的 KV-cache 管理 |
A |
沿用既有阅读记录 |
| 423 |
Google Gemma 4 12B: unified encoder-free multimodal model · 来源 2 |
A |
沿用既有阅读记录 |
| 424 |
Anthropic / Sakana AI recursive self-improvement signals |
A |
沿用既有阅读记录 |
| 425 |
Microsoft pg_durable: PostgreSQL in-database durable execution |
A |
沿用既有阅读记录 |
| 426 |
Anthropic Defending Code Reference Harness |
A |
沿用既有阅读记录 |
| 427 |
Alibaba Open Code Review: deterministic engineering + LLM agent code review |
A |
沿用既有阅读记录 |
| 428 |
Cloudflare AI Gateway spend limits / identity-driven budgets |
B |
沿用既有阅读记录 |
| 429 |
Lowfat: local CLI output filtering for agent token budgets |
B |
沿用既有阅读记录 |
| 430 |
AdaMEM: Test-Time Adaptive Memory for Language Agents · 来源 2 |
A |
沿用既有阅读记录 |
| 431 |
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents |
A |
沿用既有阅读记录 |
| 432 |
Description-Code Inconsistency in Real-world MCP Servers |
A |
沿用既有阅读记录 |
| 433 |
MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery · 来源 2 |
A |
沿用既有阅读记录 |
| 434 |
CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments |
A |
沿用既有阅读记录 |
| 435 |
Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill |
A |
沿用既有阅读记录 |
| 436 |
Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators |
A |
沿用既有阅读记录 |
| 437 |
TechCrunch: The token bill comes due |
B |
沿用既有阅读记录 |
| 438 |
Poke becomes the first AI agent on Apple Messages for Business |
B |
沿用既有阅读记录 |
| 439 |
知乎回答:王导是也缩略《置身钉内》与钉钉 ONE 项目复盘 |
B |
沿用既有阅读记录 |
| 440 |
小红书:学习如何从 0 训练一个 SOTA LLM |
B |
沿用既有阅读记录 |
| 441 |
Motus: A Unified Latent Action World Model |
A |
沿用既有阅读记录 |
| 442 |
AKO: Agentic Kernel Optimization / AKO4ALL / AKO4X · 来源 2 · 来源 3 |
S |
沿用既有阅读记录 |
| 443 |
Nano World Models: minimalist world-model experiment substrate · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 444 |
OpenAI Lockdown Mode:prompt injection 数据外泄防线的产品化 capability gate |
A |
沿用既有阅读记录 |
| 445 |
Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads |
A |
沿用既有阅读记录 |
| 446 |
SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents · 来源 2 |
A |
沿用既有阅读记录 |
| 447 |
TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration · 来源 2 |
A |
沿用既有阅读记录 |
| 448 |
Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement |
A |
沿用既有阅读记录 |
| 449 |
Google / SpaceX AI compute deal:短期 bridge capacity 与 frontier agent demand 的市场信号 |
B |
沿用既有阅读记录 |
| 450 |
Cloudflare Bot Traffic Radar:agentic web traffic 与 origin cost 进入一等指标 |
B |
沿用既有阅读记录 |
| 451 |
Tencent Productivity Agent Suite / CodeBuddy / WorkBuddy / Agent Runtime / TokenHub |
B |
沿用既有阅读记录 |
| 452 |
FrontierCode:从 correctness 到 production mergeability 的 coding-agent benchmark |
A |
沿用既有阅读记录 |
| 453 |
KV cache serving exactness:Speculative KV coding + VeriCache · 来源 2 |
A |
沿用既有阅读记录 |
| 454 |
Tokenomics / token bill:agentic software cost observability · 来源 2 |
A |
沿用既有阅读记录 |
| 455 |
Intuned Agent:browser automation codegen + managed Playwright runtime |
B |
沿用既有阅读记录 |
| 456 |
Microsoft AI developer tooling supply-chain incident |
B |
沿用既有阅读记录 |
| 457 |
AGENTS.md / context-file tooling:agent-md-bench + context file evidence |
B |
沿用既有阅读记录 |
| 458 |
SWE-Explore:repository exploration benchmark for coding agents · 来源 2 |
A |
沿用既有阅读记录 |
| 459 |
End-to-End Context Compression at Scale / LCLM · 来源 2 |
A |
沿用既有阅读记录 |
| 460 |
Anthropic biology agents:domain data infrastructure for agents |
A |
沿用既有阅读记录 |
| 461 |
Claude Fable 5 / Mythos 5:frontier capability with gated access |
B |
沿用既有阅读记录 |
| 462 |
Nemotron 3 Ultra serving stack:vLLM / SGLang / Miles day-zero path |
B |
沿用既有阅读记录 |
| 463 |
AI developer tooling security:Microsoft repo incident + SGLang RCE · 来源 2 |
B |
沿用既有阅读记录 |
| 464 |
Asuka-Bench:underspecified intent + multi-round refinement for code agents · 来源 2 |
A |
沿用既有阅读记录 |
| 465 |
DiffusionGemma:parallel text diffusion for local interactive workflows |
A |
沿用既有阅读记录 |
| 466 |
GitHub Agent Apps + Copilot Code Review skills/MCP |
A |
沿用既有阅读记录 |
| 467 |
MusaCoder:native GPU kernel generation with full-stack training on Moore Threads GPU · 来源 2 |
A |
沿用既有阅读记录 |
| 468 |
Apache Burr:state machine / telemetry / persistence for reliable AI apps |
A |
沿用既有阅读记录 |
| 469 |
Memory tools can make AI models worse:memory reliability as product risk · 来源 2 |
B |
沿用既有阅读记录 |
| 470 |
MiMo Code:long-horizon coding agent with persistent project memory · 来源 2 |
A |
沿用既有阅读记录 |
| 471 |
Claw Patrol:wire-level security firewall for agents |
A |
沿用既有阅读记录 |
| 472 |
Google DeepMind multi-agent AI safety research fund |
A |
沿用既有阅读记录 |
| 473 |
Coinbase for Agents / x402:agentic payments and paid resource access |
A |
沿用既有阅读记录 |
| 474 |
Alibaba Cloud Meoo CLI:local coding agent to cloud deployment bridge |
B |
沿用既有阅读记录 |
| 475 |
Can I Buy Your KV Cache?:agent-native prefill CDN / hot context cost model |
A |
沿用既有阅读记录 |
| 476 |
ReSum:self-summary as policy action for long reasoning RLVR |
A |
沿用既有阅读记录 |
| 477 |
AgentBeats:agentified agent assessment via A2A / MCP |
A |
沿用既有阅读记录 |
| 478 |
Agents-K1:agent-native scientific knowledge orchestration · 来源 2 · 来源 3 · 来源 4 · 来源 5 · 来源 6 |
A |
沿用既有阅读记录 |
| 479 |
EurekAgent:environment engineering for autonomous scientific discovery |
A |
沿用既有阅读记录 |
| 480 |
SkillSpector:agent skill supply-chain scanner |
A |
沿用既有阅读记录 |
| 481 |
GLM-5:from vibe coding to agentic engineering · 来源 2 · 来源 3 · 来源 4 |
A |
沿用既有阅读记录 |
| 482 |
LongTraceRL:learning long-context reasoning from search-agent trajectories · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 483 |
Plan-RewardBench / VPR:trajectory-level reward modeling and verifiable process reward for agents · 来源 2 · 来源 3 · 来源 4 · 来源 5 |
A |
沿用既有阅读记录 |
| 484 |
EvoArena / EvoMem:tracking memory evolution for robust LLM agents · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 485 |
depthfirst 21 FFmpeg zero-days:autonomous security agent with reproducible PoCs |
A |
沿用既有阅读记录 |
| 486 |
OpenAI Codex enterprise workflow cases:Notion / Nextdoor / Wasmer outcome engineering |
A |
沿用既有阅读记录 |
| 487 |
Microsoft Discovery GA:governed agentic R&D workflows |
A |
沿用既有阅读记录 |
| 488 |
TensorZero archive signal:LLMOps open-source continuity risk · 来源 2 |
B |
沿用既有阅读记录 |
| 489 |
OpenAI multistate investigation:AI product safety and personalization audit risk |
B |
沿用既有阅读记录 |
| 490 |
WeaveBench:hybrid-interface long-horizon computer-use agent benchmark · 来源 2 |
A |
沿用既有阅读记录 |
| 491 |
TRACE:compiling user corrections into runtime enforcement for coding agents · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 492 |
HarnessBridge:learnable bidirectional controller for LLM agent harness · 来源 2 |
A |
沿用既有阅读记录 |
| 493 |
EvoBrowseComp:benchmarking search agents on evolving knowledge · 来源 2 |
A |
沿用既有阅读记录 |
| 494 |
Perplexity / HBS:How AI Agents Reshape Knowledge Work |
A |
沿用既有阅读记录 |
| 495 |
Google DeepMind From AGI to ASI:multi-agent collectives as one ASI pathway · 来源 2 |
B |
沿用既有阅读记录 |
| 496 |
OpenAI Partner Network:enterprise AI delivery and specialization market |
B |
沿用既有阅读记录 |
| 497 |
Context window budget:Don't trust large context windows |
B |
沿用既有阅读记录 |
| 498 |
AI provenance failure signal:UK police fake-evidence allegation + KPMG hallucinated report |
B |
沿用既有阅读记录 |
| 499 |
Gabriel Weinberg:No, everyone is not using AI for everything |
B |
沿用既有阅读记录 |
| 500 |
HarnessX:composable, adaptive, evolvable agent harness foundry |
A |
沿用既有阅读记录 |
| 501 |
StreamMemBench:streaming evaluation of agent memory for future-oriented assistance · 来源 2 |
A |
沿用既有阅读记录 |
| 502 |
Dialogue SWE-Bench:benchmarking dialogue-driven coding agents · 来源 2 |
A |
沿用既有阅读记录 |
| 503 |
Parallel-Synthesis:direct latent-space synthesis for parallel branches in LLM-agent workflows |
A |
沿用既有阅读记录 |
| 504 |
SIMMER:latent failures in LLM executable planning with a world model · 来源 2 |
A |
沿用既有阅读记录 |
| 505 |
Google Cloud OKF:Open Knowledge Format for agent-readable context bundles · 来源 2 |
A |
沿用既有阅读记录 |
| 506 |
OpenRouter Fusion:model panels as an API-level reasoning primitive |
A |
沿用既有阅读记录 |
| 507 |
Apple Foundation Models / Xcode 27:native provider protocol and agentic coding workflow |
A |
沿用既有阅读记录 |
| 508 |
Enterprise agent identity and service-agent consolidation:NewCore + Salesforce / Fin |
B |
沿用既有阅读记录 |
| 509 |
Measuring Agents in Production:真实生产 agent 仍是高频人工介入系统 |
A |
沿用既有阅读记录 |
| 510 |
Principles of Mixed-Initiative User Interfaces:混合主动权的经典设计原则 |
A |
沿用既有阅读记录 |
| 511 |
Power to the People:Interactive ML 中人的角色 |
A |
沿用既有阅读记录 |
| 512 |
Evaluation of Interactive Machine Learning Systems:algorithm-centered + human-centered 双验证 |
A |
沿用既有阅读记录 |
| 513 |
A Benchmark for Scalable Oversight Mechanisms:监督机制也需要 benchmark |
A |
沿用既有阅读记录 |
| 514 |
Deep Reinforcement Learning from Human Preferences:少量偏好监督如何塑造复杂目标 |
A |
沿用既有阅读记录 |
| 515 |
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments |
S |
沿用既有阅读记录 |
| 516 |
Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents |
S |
沿用既有阅读记录 |
| 517 |
OpenRouter Fusion:多模型 deliberation 作为 API runtime primitive |
A |
沿用既有阅读记录 |
| 518 |
OpenAI Deployment Simulation:用真实分布预演模型上线行为 |
S |
沿用既有阅读记录 |
| 519 |
AA-AgentPerf:agentic inference 的 SLO / agents-per-megawatt 口径 |
A |
沿用既有阅读记录 |
| 520 |
GLM-5.2:open-weight long-horizon agent model · 来源 2 |
A |
沿用既有阅读记录 |
| 521 |
AI Coding Agents Can Reproduce Social Science Findings / SocSci-Repro-Bench |
A |
沿用既有阅读记录 |
| 522 |
LifeSciBench + AI Chemist:science agent 的专家 rubric 与湿实验闭环 |
A |
沿用既有阅读记录 |
| 523 |
Appia / Pramaana:AI trust 从 policy 走向 conformity + proof |
A |
沿用既有阅读记录 |
| 524 |
MCP Enterprise-Managed Authorization:Zero-touch OAuth for MCP |
A |
沿用既有阅读记录 |
| 525 |
Elastic agent memory:hybrid retrieval + DLS 的生产 memory 参考 |
A |
沿用既有阅读记录 |
| 526 |
Decoupled Search Grounding:把 agent search 变成 MCP-compatible gateway |
A |
仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 |
| 527 |
RODS:multi-turn tool-use agent 的 reward-driven online data synthesis |
A |
仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 |
| 528 |
CEO-Bench:long-horizon business agent eval |
A |
仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 |
| 529 |
SGCD:GUI agent 的 off-trajectory continuation distillation |
A |
仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 |
| 530 |
Xcientist:AI scientist 的 research harness 与 claim drift 防线 |
A |
仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 |
| 531 |
Stanford PhD 回山东做传统企业 AI 落地:代码快,诊断慢 |
A |
沿用既有阅读记录 |
| 532 |
Google Agentic Resource Discovery:agent capability discovery + trust manifest |
A |
沿用既有阅读记录 |
| 533 |
Claude Design + Claude Code sync:design-to-code workspace 进入真实组件回路 |
A |
沿用既有阅读记录 |
| 534 |
SkillVetBench:LLM agent skills 的语义风险评估 |
A |
沿用既有阅读记录 |
| 535 |
ADK Arena:用 LLM-as-a-Developer 测 agent framework 可用性 |
A |
沿用既有阅读记录 |
| 536 |
Agent Planning Benchmark:把 planning failure 从执行失败里拆出来 |
A |
沿用既有阅读记录 |
| 537 |
R3-Skill:skill routing 中的 rejection signal 不是垃圾数据 |
A |
沿用既有阅读记录 |
| 538 |
Exploration Structure in LLM Agents:coding agent 的 repo traversal 结构会决定定位质量 |
A |
沿用既有阅读记录 |
| 539 |
AgentFairBench:agent 公平性要测 action,不只测 answer |
A |
沿用既有阅读记录 |
| 540 |
Kimi Work / Kimi Code Goal Mode:本地长程 agent 的 goal state 与权限面 |
A |
沿用既有阅读记录 |
| 541 |
OpenAI Patch the Planet:安全 agent 的 patch loop 与 maintainer agency |
A |
沿用既有阅读记录 |
| 542 |
Google Interactions API GA:managed agent API contract |
A |
沿用既有阅读记录 |
| 543 |
Multi-LCB + Contagion Networks:coding agent 评测的语言轴与 judge topology · 来源 2 |
A |
沿用既有阅读记录 |
| 544 |
Liquid AI LFM2.5 Retrievers:本地多语言 memory/search 检索底座 |
A |
沿用既有阅读记录 |
| 545 |
xAI Grok Build /goal:long-running coding agent 的目标状态与验证面 |
A |
沿用既有阅读记录 |
| 546 |
Claude Tag:Slack 中的 scoped team agent 与组织级权限/成本控制 |
A |
沿用既有阅读记录 |
| 547 |
The Coming Loop:harness-level loop 会放大工程质量债 |
A |
沿用既有阅读记录 |
| 548 |
Mistral OCR 4 + Baidu Unlimited-OCR:文档 ingestion 从 OCR 走向结构化 context substrate · 来源 2 |
A |
沿用既有阅读记录 |
| 549 |
Randomized YaRN:短上下文训练也能改善 16K-128K 长上下文推理泛化 |
A |
沿用既有阅读记录 |
| 550 |
AIR:用 RL 学会何时在多模态推理中调用代码工具 |
A |
沿用既有阅读记录 |
| 551 |
Can LLMs Reliably Self-Report Adversarial Prefills:模型自我报告不能当安全证据 |
A |
沿用既有阅读记录 |
| 552 |
VibeThinker-3B:小模型 verifiable reasoning 的成本/能力边界 |
A |
沿用既有阅读记录 |
| 553 |
Gemini 3.5 Flash computer use:UI agent 的 observe-act-screenshot contract |
A |
沿用既有阅读记录 |
| 554 |
Qwen-AgentWorld:language world model for agentic RL · 来源 2 |
A |
沿用既有阅读记录 |
| 555 |
AOHP:Android Open Harness Project / OS-level agent harness · 来源 2 |
A |
沿用既有阅读记录 |
| 556 |
OpenThoughts-Agent:agentic model data recipes · 来源 2 |
A |
沿用既有阅读记录 |
| 557 |
Tmax:terminal-agent RL recipe and TMax-15K · 来源 2 · 来源 3 |
A |
沿用既有阅读记录 |
| 558 |
OpenAI/Broadcom Jalapeno:inference cost as agent runtime constraint |
A |
沿用既有阅读记录 |
| 559 |
Qualcomm 收购 Modular:portable AI serving/software stack signal |
A |
沿用既有阅读记录 |
| 560 |
Headroom:agent context compression / CCR gate |
A |
沿用既有阅读记录 |
| 561 |
Notion Mail agent takeover + General Intuition action-labeled world model data |
A |
沿用既有阅读记录 |
| 562 |
OpenAI GPT-5.6 Sol limited preview:frontier model release gate and ultra subagents |
A |
沿用既有阅读记录 |
| 563 |
OpenAI Codex economic research:agents transform work into delegated long-horizon tasks |
A |
沿用既有阅读记录 |
| 564 |
Cursor reward hacking in coding benchmarks:strict harness for aware coding agents |
A |
沿用既有阅读记录 |
| 565 |
Workweave Router:cache-aware model routing for Claude Code / Codex / Cursor |
A |
沿用既有阅读记录 |
| 566 |
NVIDIA NeMo AutoModel / Transformers v5 Expert Parallelism + DeepEP MoE fine-tuning path · 来源 2 |
A |
沿用既有阅读记录 |
| 567 |
PEEU GUI agents:Autonomous Experience Exploration + Hindsight Experience Utilization for task planning |
A |
沿用既有阅读记录 |
| 568 |
When are likely answers right? Sequence Probability and Correctness in LLMs |
A |
沿用既有阅读记录 |
| 569 |
Un-0:open coupled-oscillator image generator as physical-compute substrate · 来源 2 |
A |
沿用既有阅读记录 |
| 570 |
OpenAI + Broadcom Jalapeno inference chip:full-stack LLM inference platform |
A |
沿用既有阅读记录 |
| 571 |
DeepSeek DSpark / DeepSpec:confidence-scheduled speculative decoding for production serving · 来源 2 |
A |
沿用既有阅读记录 |
| 572 |
AWS Lambda MicroVMs:full-lifecycle Firecracker sandboxes for AI agents |
A |
沿用既有阅读记录 |
| 573 |
DBOSify:Postgres-backed durable workflow as compact Temporal alternative |
A |
沿用既有阅读记录 |
| 574 |
Adrafinil:macOS activity assertion layer for long-running AI agents |
A |
沿用既有阅读记录 |
| 575 |
Cloud World Model:cloud-infra simulation product radar |
A |
仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 |
| 576 |
Are We Ready For An Agent-Native Memory System? Trustworthy Memory Search · 来源 2 |
A |
沿用既有阅读记录 |
| 577 |
Semgrep GLM-5.2 cyber benchmark:real security harness for coding agents |
A |
沿用既有阅读记录 |
| 578 |
GitHub Copilot App BYOK:provider/account boundary for agent sessions |
A |
沿用既有阅读记录 |
| 579 |
Wayfinder Router:deterministic local/cloud LLM routing |
A |
沿用既有阅读记录 |
| 580 |
OpenAI Codex issue:exclude sensitive files by default |
A |
沿用既有阅读记录 |
| 581 |
Tokenmaxxing / agent output budget policy |
A |
沿用既有阅读记录 |
| 582 |
百度千帆 Coding Plan -> Token Plan:coding agent usage ledger signal |
B |
沿用既有阅读记录 |
| 583 |
Dual Strix Halo vLLM cluster:local serving lab reference |
B |
沿用既有阅读记录 |
| 584 |
Claude Sonnet 5:agent default model cost-performance reset |
A |
沿用既有阅读记录 |
| 585 |
Claude Science AI workbench:auditable domain agent OS |
A |
沿用既有阅读记录 |
| 586 |
Gemini Spark updates:desktop automation + connected apps + custom MCP |
A |
沿用既有阅读记录 |
| 587 |
vLLM Micro-Agent:serving router as bounded agent collaboration |
A |
沿用既有阅读记录 |
| 588 |
SWE-Together:interactive user-session benchmark for coding agents |
A |
沿用既有阅读记录 |
| 589 |
OSWorld 2.0:long-horizon computer-use official runner boundary |
A |
沿用既有阅读记录 |
| 590 |
SWE-MeM:adaptive memory management for long-horizon coding agents |
A |
沿用既有阅读记录 |
| 591 |
Language Firewall:routing defense for multi-agent systems |
A |
沿用既有阅读记录 |
| 592 |
Couchbase AI Data Plane:enterprise memory/context substrate |
A |
沿用既有阅读记录 |
| 593 |
Lingtai:local-first lifelong Agent runtime |
A |
沿用既有阅读记录 |
| 594 |
Ephemeral Sandbox:COW workspace 与 OCC publication |
A |
沿用既有阅读记录 |
| 595 |
Using Claude Code: The Unreasonable Effectiveness of HTML |
A |
沿用既有阅读记录 |
| 596 |
梁文锋投资者交流会 · 网传录音文字稿(非官方) · 来源 2 |
A |
仅元信息;已取得 42 页 PDF、提取文本并检查首页,未逐段精读。网传转写稿,未经 DeepSeek 或本人公开确认;ASR 与 AI 整理可能引入错误,不作为官方口径。 |
| 597 |
hai-stack / Geju:用 target model、falsifier 与收益账单抵抗局部补丁化 |
A |
沿用既有阅读记录 |
| 598 |
Waza:把工程习惯产品化为可路由、可验证的 Agent skills |
A |
沿用既有阅读记录 |
| 599 |
BfdCampos Mermaid skill:把“语法正确”提升为“渲染后可读” |
A |
沿用既有阅读记录 |
| 600 |
Oracle:把第二模型意见封装成可追踪的 consult session |
A |
沿用既有阅读记录 |
| 601 |
Graph Engineering:给 Agent Harness 叠加显式控制流图,而不是换一个新名词 |
A |
沿用既有阅读记录 |
| 602 |
赵克常《炒股挣钱》:风险认知课,不是可复制的投资研究方法 |
B |
沿用既有阅读记录 |
| 603 |
未核实的 MLSys 截图线索:集中式推理、确定性信号与知识图谱控制面 |
B |
沿用既有阅读记录;来源真实性未确认;保留为核验案例,不作为论文结论。 |
| 604 |
Evolvent AI GitHub 组织复核:与既有研究目录的重复记录 |
B |
沿用既有阅读记录;保留稳定 ID 与重复记录;不因重复而静默删除。 |