Agent architecture
Structural design of agent systems: planning loops, sub-agent decomposition, state management, and control flow.
tagged in 1 of 97 digest issues, most recently 2026-06-22
Agent memory
How agents store, retrieve, and forget context across turns and sessions; memory architectures and their design tradeoffs.
tagged in 50 of 97 digest issues, most recently 2026-07-22
Digest issues
- The model is the constant; the harness is the variable — 2026-07-22
- 96 million tokens saved on paper, a 7.6% higher bill in practice — 2026-07-21
- A rewritten agent loop beat maximum reasoning effort at 40% of the cost — 2026-07-20
- 96 million tokens "saved" and a 7.6% higher bill — 2026-07-20
- The Bridge Documents Your Reranker Buries — 2026-07-19
- Coding-Agent Gains Are Moving From the Model to the Harness — 2026-07-18
- An Agent Proved TLS 1.3 Correct. It Would Also Install Your Malware From a Bad README. — 2026-07-17
- Harness evolution doesn't beat plain test-time scaling on Terminal-Bench 2.1 — 2026-07-16
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- Claude Cowork moves to the cloud, and 90% of its sessions aren't coding — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- AI pull requests carry 1.7× more defects, and review didn't scale to match — 2026-06-27
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Anthropic Puts Claude in Slack — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Enhancing Developer Productivity with Google Colab CLI and Agentic Observability — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
Agent observability
Tracing, monitoring, and debugging agent runs in production: telemetry, replay, and failure analysis.
tagged in 1 of 97 digest issues, most recently 2026-06-25
Agent tooling
Harnesses, skills, and tool interfaces that let agents act: the plumbing between a model and the systems it operates.
tagged in 44 of 97 digest issues, most recently 2026-07-22
Digest issues
- Kimi K3 clears the open-weight bar, and an OpenAI cyber model clears the sandbox — 2026-07-22
- A Jacobian Conjecture counterexample, and METR's worst cheating rate yet — 2026-07-20
- Kimi K3 arrives at 2.8T parameters, and everyone starts routing by price tier — 2026-07-20
- Bun-in-Rust ships silently, Gemini slips loudly — 2026-07-19
- Fable 5 Stays, and the Benchmarks Get a Harder Look — 2026-07-18
- Kimi K3 Prices Like Sonnet, and Everyone Wants the Agent's Terminal — 2026-07-17
- Thinking Machines ships Inkling: 975B parameters, Apache 2.0, and 25K tokens per task — 2026-07-16
- Frontier agents ace one task and stall across a stream of them — 2026-07-13
- OpenAI claims a math proof, walks back its launch, and the model race turns to cost — 2026-07-12
- Claude Cowork moves to the cloud, and 90% of its sessions aren't coding — 2026-07-11
- GPT-5.6 lands, and the coding-model price floor drops again — 2026-07-10
- Grok 4.5 Ships as Cursor's First General-Purpose Model, and OpenAI Retracts SWE-Bench Pro — 2026-07-09
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- x402 payments reach both edge networks, and GPT-5.6 Sol Ultra heads to Codex — 2026-07-06
- Sonnet 5 Lands, Fable 5 Returns, and ZCode Undercuts Everyone on Price — 2026-07-06
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Fable 5 completes the task as a test and refuses it as production work — 2026-07-04
- The field can write more code than it can understand — 2026-07-03
- Claude Tag lands 65% of Anthropic's internal PRs while the World's Fair argues over the outer loop — 2026-07-02
- Claude Sonnet 5 lands everywhere at once — 2026-07-01
- The frontier model becomes the supervisor — 2026-06-30
- Open weights became the default coding model the week the frontier got restricted — 2026-06-29
- The floor is rising: open models, the coding-speed paradox, and the bill for the buildout — 2026-06-29
- GPT-5.6 and Mythos 5 ship behind a government access gate — 2026-06-28
- OpenAI ships GPT-5.6 behind a government gate, the same day Washington un-blocks Mythos 5 — 2026-06-27
- Codex usage at OpenAI jumped 56x, and the agent stack started grading itself — 2026-06-26
- Open weights match Opus at half the cost, and ship free in Devin — 2026-06-25
- Anthropic Puts Claude in Slack — 2026-06-24
- Cyber models ship faster than the rules for them — 2026-06-23
- GLM 5.2 edges past Sonnet on 1,000 coding tasks as Claude's error rates spike — 2026-06-22
- GLM-5.2 takes the open-model crown while agents settle into production — 2026-06-22
- A single page can RCE your agent's host — 2026-06-21
- Anthropic resets every usage limit while negotiating Fable back from a US ban — 2026-06-20
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- Two labs put their models on the lab bench — 2026-06-18
- GLM-5.2 cracks the open-weight coding frontier, and Cursor goes to SpaceX — 2026-06-17
- The Fable 5 "jailbreak" was "fix this code" — 2026-06-16
- When the loop closes, verification becomes the job — 2026-06-15
- Anthropic shipped the best coding model measured, then the government pulled it — 2026-06-15
- The model layer becomes a regulated surface — 2026-06-14
- The day the US government switched off Fable 5 — 2026-06-13
- OpenAI buys Ona while Fable 5 starts downgrading itself — 2026-06-12
- Fable 5 scores 91, real code scores 13 — 2026-06-11
- Claude Fable 5 arrives at twice Opus pricing, and Cognition's day-old FrontierCode crowns it #1 — 2026-06-10
Agentic coding
Agents that write, refactor, and maintain software: coding assistants, autonomous dev loops, and the workflows around them.
tagged in 60 of 97 digest issues, most recently 2026-07-22
Digest issues
- The model is the constant; the harness is the variable — 2026-07-22
- Kimi K3 clears the open-weight bar, and an OpenAI cyber model clears the sandbox — 2026-07-22
- 96 million tokens saved on paper, a 7.6% higher bill in practice — 2026-07-21
- A rewritten agent loop beat maximum reasoning effort at 40% of the cost — 2026-07-20
- 96 million tokens "saved" and a 7.6% higher bill — 2026-07-20
- The Bridge Documents Your Reranker Buries — 2026-07-19
- Coding-Agent Gains Are Moving From the Model to the Harness — 2026-07-18
- An Agent Proved TLS 1.3 Correct. It Would Also Install Your Malware From a Bad README. — 2026-07-17
- Harness evolution doesn't beat plain test-time scaling on Terminal-Bench 2.1 — 2026-07-16
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- Grok Build CLI uploaded whole repos to a Google bucket; Codex hit 7M users — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- Frontier agents ace one task and stall across a stream of them — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Sonnet 5 Lands, Fable 5 Returns, and ZCode Undercuts Everyone on Price — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The field can write more code than it can understand — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- The floor is rising: open models, the coding-speed paradox, and the bill for the buildout — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- AI pull requests carry 1.7× more defects, and review didn't scale to match — 2026-06-27
- The agent scaffold is now the contested layer — 2026-06-26
- Codex usage at OpenAI jumped 56x, and the agent stack started grading itself — 2026-06-26
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
AI agents
Systems that plan, call tools, and act over multiple steps to accomplish a goal.
tagged in 0 of 97 digest issues
AI and the labor market
AI's effect on jobs and work: displacement, augmentation, and workforce shifts.
tagged in 3 of 97 digest issues, most recently 2026-07-03
AI economics
Cost structure of AI: token pricing, inference margins, usage limits, and unit economics.
tagged in 17 of 97 digest issues, most recently 2026-07-22
Digest issues
- Kimi K3 clears the open-weight bar, and an OpenAI cyber model clears the sandbox — 2026-07-22
- A Jacobian Conjecture counterexample, and METR's worst cheating rate yet — 2026-07-20
- Kimi K3 arrives at 2.8T parameters, and everyone starts routing by price tier — 2026-07-20
- Bun-in-Rust ships silently, Gemini slips loudly — 2026-07-19
- OpenAI claims a math proof, walks back its launch, and the model race turns to cost — 2026-07-12
- GPT-5.6 lands, and the coding-model price floor drops again — 2026-07-10
- x402 payments reach both edge networks, and GPT-5.6 Sol Ultra heads to Codex — 2026-07-06
- Fable 5 completes the task as a test and refuses it as production work — 2026-07-04
- Claude Sonnet 5 lands everywhere at once — 2026-07-01
- OpenAI ships GPT-5.6 behind a government gate, the same day Washington un-blocks Mythos 5 — 2026-06-27
- Open weights match Opus at half the cost, and ship free in Devin — 2026-06-25
- A single page can RCE your agent's host — 2026-06-21
- Anthropic resets every usage limit while negotiating Fable back from a US ban — 2026-06-20
- GLM-5.2 cracks the open-weight coding frontier, and Cursor goes to SpaceX — 2026-06-17
- The Fable 5 "jailbreak" was "fix this code" — 2026-06-16
- Fable 5 scores 91, real code scores 13 — 2026-06-11
- Claude Fable 5 arrives at twice Opus pricing, and Cognition's day-old FrontierCode crowns it #1 — 2026-06-10
AI for science
AI applied to scientific discovery and research workflows, from literature synthesis to hypothesis generation.
tagged in 7 of 97 digest issues, most recently 2026-07-14
AI governance
Rules for AI systems: regulation, policy, export controls, and organizational governance of agents.
tagged in 13 of 97 digest issues, most recently 2026-07-13
AI industry
The business landscape of AI: labs, funding, acquisitions, and competitive strategy.
tagged in 5 of 97 digest issues, most recently 2026-07-14
AI infrastructure
Compute, serving, and platform layers under AI systems: GPUs, inference stacks, and datacenter buildout.
tagged in 9 of 97 digest issues, most recently 2026-06-29
AI safety
Preventing harmful model and agent behavior: alignment, oversight, and safety evaluation.
tagged in 10 of 97 digest issues, most recently 2026-07-16
AI security
Securing AI systems and using AI in security: prompt injection, sandboxing, supply-chain risk, and agent attack surfaces.
tagged in 24 of 97 digest issues, most recently 2026-07-22 · 18 papers · 1 explorer section
Digest issues
- Kimi K3 clears the open-weight bar, and an OpenAI cyber model clears the sandbox — 2026-07-22
- Kimi K3 arrives at 2.8T parameters, and everyone starts routing by price tier — 2026-07-20
- Fable 5 Stays, and the Benchmarks Get a Harder Look — 2026-07-18
- An Agent Proved TLS 1.3 Correct. It Would Also Install Your Malware From a Bad README. — 2026-07-17
- Kimi K3 Prices Like Sonnet, and Everyone Wants the Agent's Terminal — 2026-07-17
- Thinking Machines ships Inkling: 975B parameters, Apache 2.0, and 25K tokens per task — 2026-07-16
- Grok Build CLI uploaded whole repos to a Google bucket; Codex hit 7M users — 2026-07-14
- OpenAI claims a math proof, walks back its launch, and the model race turns to cost — 2026-07-12
- Claude Cowork moves to the cloud, and 90% of its sessions aren't coding — 2026-07-11
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Fable 5 completes the task as a test and refuses it as production work — 2026-07-04
- The field can write more code than it can understand — 2026-07-03
- Open weights became the default coding model the week the frontier got restricted — 2026-06-29
- The floor is rising: open models, the coding-speed paradox, and the bill for the buildout — 2026-06-29
- Open weights match Opus at half the cost, and ship free in Devin — 2026-06-25
- Anthropic Puts Claude in Slack — 2026-06-24
- Cyber models ship faster than the rules for them — 2026-06-23
- A single page can RCE your agent's host — 2026-06-21
- Two labs put their models on the lab bench — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- When the loop closes, verification becomes the job — 2026-06-15
- The model layer becomes a regulated surface — 2026-06-14
- The day the US government switched off Fable 5 — 2026-06-13
Papers
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Attackers compromise LLM-integrated apps by planting instructions in content the model later retrieves (web pages, documents) — no direct access needed.
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents — A dynamic benchmark of realistic tool-using tasks for measuring prompt-injection attacks and defenses — and showing current defenses are far from complete.
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-Agent Systems — Detects anomalies over the agent communication graph and intervenes to defend a multi-agent system against attacks that propagate between agents.
- Red-Teaming LLM Multi-Agent Systems via Communication Attacks — Injects malicious messages that propagate through inter-agent communication channels, compromising a multi-agent system from the inside.
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions — Surveys the security threat landscape of MCP — the tool-connection standard now powering most agentic systems — and its expanding attack surface.
- MPAC: A Multi-Principal Agent Coordination Protocol (extends MCP + A2A)
- Architecture Matters for Multi-Agent Security — The same task under different multi-agent architectures has different security properties.
- AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents — Plants deception traps (fake tools/credentials) that detect indirect-prompt-injection compromises which slip past prevention — cross-lingual, near-zero false positives.
- Exploiting LLM Agent Supply Chains via Payload-Less Skills — Semantic Compliance Hijacking (SCH) weaponizes third-party Agent Skill files without any embedded code: it injects obfuscated natural-language 'compliance guidelines' that the agent's retrieve-contextualize-execute pipeline fuses directly into the LLM system prompt as authoritative directives, and the agent then synthesizes and runs the malicious logic itself at runtime (dumping os.environ and POSTing it as 'telemetry', or exec()-ing a fetched remote byte stream). Across 600 contextualized tasks over 3 frameworks (OpenClaw, Claude Code, Codex) and 3 models, SCH hit a peak 77.67% complete-leakage rate and 67.33% RCE, yet both SkillScan (AST/regex) and LLM Guard (semantic classifiers) flagged it at 0.00%, versus 62.00-96.00% detection for SkillJect and 91.61-99.81% for DDIPE explicit-payload baselines; an 'Alignment-Security Paradox' made the most-aligned model (GPT-5.4 mini, 62.11% avg complete leakage) the most exploitable because it best maps prose directives to concrete system APIs.
- AI Assurance: A Comprehensive Testing Strategy for Enterprise AI Systems
- Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents — PORTICO is a reference monitor that gives coding-agent tools a bounded capability lifetime: an explicit task contract compiles into an initial envelope plus grant, trusted-closure, and global-deny rules, and a request-grant-invoke protocol mints opaque, epoch-bound handles that are revoked when the subgoal closes, so authority cannot linger or be replayed once its justifying episode ends.
- AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent — A white-box runtime defense for multi-turn coding agents that drive a shell in Docker. At each step it reads the last-token hidden state at one layer and runs a single linear probe to flag a harmful execution state; on a flag it steers the representation inside a sparse 10-dimensional subspace, adapting only the steering strength via an LLM judge scoring safety (0.6) and utility (0.4). It detects current-step risk at 97.32% average accuracy, predicts the next harmful command at up to 96.77% lookahead, and cuts average attack success from 85.99% to 13.36% (72.63 pp), beating RepE and self-reminder. Negative steering of the same subspace flips genuine refusals into malicious commands (100% ASR on LLaMA), confirming the direction is causal not lexical; under prompt injection the probe misfires but applying the steering direction still drops ASR 86.7%->6.7%.
- Knowledge-Based Pull Requests: A Trusted Workflow for Agent-Mediated Knowledge Collaboration — KPR treats an external collaborator's code, tests, and agent trace as a knowledge source, not a merge candidate; only a project-owned inner trusted coding agent regenerates code inside the receiving project's environment, under repository context and security policy, before anything can be merged.
- From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents — Governance holds only because it is code-owned, not prompted. With the model fixed and only the enforcement layer varied, prompt instructions alone let recommendation-language and internal-trace-leakage violations reach the reader on all 30 adversarial runs, every one of which the harness blocked; only source-backed claims are allowed to enter runtime context in the first place. Across 270 composition-boundary runs under three substituted models the contracts held, with failures confined to the model-composed side and recorded.
- ACRFence: Preventing Semantic Rollback Attacks in Agent Checkpoint-Restore
- AgentSight: System-Level Observability for AI Agents Using eBPF
- NIST AI Risk Management Framework (AI RMF 1.0)
- OWASP Top 10 for LLM Applications — A community-maintained catalogue of the top LLM-application security risks — prompt injection, insecure output handling, excessive agency, supply chain, and more.
Code intelligence
Understanding codebases at scale: search, navigation, and agents that reason over source.
tagged in 0 of 97 digest issues
Code review
Automated and agent-assisted review of code changes: quality gates, review bots, and human-agent review workflows.
tagged in 7 of 97 digest issues, most recently 2026-07-13
Computer use
Agents that operate GUIs and browsers directly: screen perception, action models, and their reliability limits.
tagged in 1 of 97 digest issues, most recently 2026-06-25
Context engineering
Deciding what goes into a model's context window and when: packing, pruning, and structuring working context for long-running tasks.
tagged in 5 of 97 digest issues, most recently 2026-07-18
Developer productivity
How AI tooling changes software work: measured impact, adoption patterns, and workflow shifts.
tagged in 6 of 97 digest issues, most recently 2026-07-16
Enterprise adoption
How organizations deploy AI in production: procurement, integration, and the gap between demos and durable systems.
tagged in 2 of 97 digest issues, most recently 2026-06-22
Evaluation
Measuring whether AI systems work: benchmark design, eval harnesses, and comparison of models, agents, and search systems.
tagged in 65 of 97 digest issues, most recently 2026-07-22 · 19 papers · 2 explorer sections
Digest issues
- The model is the constant; the harness is the variable — 2026-07-22
- Kimi K3 clears the open-weight bar, and an OpenAI cyber model clears the sandbox — 2026-07-22
- 96 million tokens saved on paper, a 7.6% higher bill in practice — 2026-07-21
- A rewritten agent loop beat maximum reasoning effort at 40% of the cost — 2026-07-20
- A Jacobian Conjecture counterexample, and METR's worst cheating rate yet — 2026-07-20
- 96 million tokens "saved" and a 7.6% higher bill — 2026-07-20
- Kimi K3 arrives at 2.8T parameters, and everyone starts routing by price tier — 2026-07-20
- The Bridge Documents Your Reranker Buries — 2026-07-19
- Coding-Agent Gains Are Moving From the Model to the Harness — 2026-07-18
- Fable 5 Stays, and the Benchmarks Get a Harder Look — 2026-07-18
- An Agent Proved TLS 1.3 Correct. It Would Also Install Your Malware From a Bad README. — 2026-07-17
- Harness evolution doesn't beat plain test-time scaling on Terminal-Bench 2.1 — 2026-07-16
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- Grok 4.5 Ships as Cursor's First General-Purpose Model, and OpenAI Retracts SWE-Bench Pro — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Claude Tag lands 65% of Anthropic's internal PRs while the World's Fair argues over the outer loop — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- The frontier model becomes the supervisor — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- AI pull requests carry 1.7× more defects, and review didn't scale to match — 2026-06-27
- The agent scaffold is now the contested layer — 2026-06-26
- Codex usage at OpenAI jumped 56x, and the agent stack started grading itself — 2026-06-26
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- GLM 5.2 edges past Sonnet on 1,000 coding tasks as Claude's error rates spike — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Fable 5 scores 91, real code scores 13 — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- Claude Fable 5 arrives at twice Opus pricing, and Cognition's day-old FrontierCode crowns it #1 — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Enhancing Developer Productivity with Google Colab CLI and Agentic Observability — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
Papers
- Evaluating Very Long-Term Conversational Memory of LLM Agents — Introduces LoCoMo, the canonical long-term conversational-memory benchmark: very-long dyadic (person-person) dialogues spanning up to ~35 sessions / hundreds of turns, generated via a machine-human pipeline with personas and temporal event graphs, plus a QA test suite, event-summarization, and multimodal-dialogue-generation tasks.
- Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations — Surveys 22 personalized-dialogue datasets and 17 seminal works (2021-2023), cataloging persona/personalization datasets, problem types, and evaluation facets/metrics for personalized dialogue.
- Agent Workflow Memory — Introduces Agent Workflow Memory, inducing reusable workflows from past trajectories to improve long-horizon web-navigation agents, evaluated on agent-trajectory benchmarks (e.g. web tasks) rather than conversational QA.
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Defines LongMemEval, a chat-assistant memory benchmark organized around five core abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention), with controllable interaction histories so memory load can be scaled independently of the target evidence.
- REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation — Releases REALTALK, 21 days of genuine human-human messaging conversations (not LLM-synthesized), used to study long-term memory probing and persona/emotional-intelligence simulation against the distribution gap between synthetic and real dialogue.
- Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey — PRISMA-based survey of evaluation methods for LLM agents in multi-turn conversational settings, cataloging benchmarks, metrics, and methodologies across the multi-turn evaluation landscape.
- ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models — Proposes ImplicitMemBench to measure implicit (procedural) memory, where past experience becomes automated behavior, rather than the explicit fact-recall that existing memory benchmarks test.
- Evaluating Memory Capability in Continuous Lifelog Scenario — Argues existing memory benchmarks cover only Person-AI and dyadic Person-Person dialogue and neglects 'continuous dialogue lifelogs' from always-on wearables; introduces lifelog memory evaluation resources (LifeMem / EgoMem first-person scenarios) for ambient continuously-recorded conversation.
- Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo — Pins down LoCoMo's exact composition (ACL 2024 version: 10 conversations, 1,813 questions) and argues every LoCoMo system treats memory as retrieval over dialogue segments while none reports adversarial robustness (refusing questions about facts never disclosed); proposes that abstention/hallucination-resistance metric as a missing axis.
- AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs — Presents a topic-guided multi-agent conversation-generation pipeline that synthesizes dialogues specifically for fine-tuning and evaluating both short- and long-term memory.
- GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations — Introduces GroupMemBench, targeting memory in multi-party (group) conversations where the agent must attribute facts to specific speakers across many participants, a setting that dyadic benchmarks like LoCoMo/LongMemEval do not cover.
- MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models — Introduces MemLens to benchmark long-term memory in vision-language models across long multimodal interactions, evaluating whether memory methods preserve evidence needed for later recall.
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory — Proposes MemEye, a visual-centric evaluation framework that specifically tests whether agents preserve the visual evidence required for later recall in multimodal memory.
- MemFail: Stress-Testing Failure Modes of LLM Memory Systems — Builds a stress-test benchmark that systematically enumerates and probes failure modes of external memory systems (consistency breaks over long-horizon interactions) rather than aggregate accuracy.
- Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison — Constructs a diagnostic testbed for selective QA over conflicting multi-source personal memory, where the system must decide how to use contradictory stored facts, plus a comparison of methods on that testbed.
- GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge — Separates the value of a substrate from raw accuracy: per-backend cost is measured, and a similarity sweep shows memory pays only above a 'copyability threshold' — when the retrieved case is a near-duplicate (cosine ≳ 0.8) accuracy jumps +12 to +13.5 pp, while below it nothing helps; the gain is answer retrieval, not method transfer, and a 4.5× larger backbone steepens the near-duplicate step (+22.5 to +28.5 pp) yet still cannot extract a transferable method from a worked example.
- StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance — A streaming benchmark sourced from real EgoLife egocentric lifelogs rather than scripted or synthesized dialogue: each five-minute segment is mined for a hidden 'evidence anchor' (a user-specific preference, plan, or capability) that spawns a two-step task sequence — an initial task that needs the evidence and a follow-up task grounded in the same anchor; eight memory systems are evaluated across two backbones (DeepSeek-V4-Flash, Gemini-3-Flash).
- MemTrace: Probing What Final Accuracy Misses in Long-Term Memory — MemTrace contains 835 typed knowledge points drawn from 20 users, expanded into 15,422 question rows and over 200,000 scored answers, and is used to evaluate 13 memory-system configurations spanning four paradigms: long-context models, retrieval-augmented systems, external-memory stores, and agentic-memory architectures.
- MemSyco-Bench: Benchmarking Sycophancy in Agent Memory — MemSyco-Bench covers five task categories testing whether agents can reject memory as factual evidence, respect its scope, resolve memory-vs-evidence conflicts, track updates, and use valid memory for personalization, built after showing existing benchmarks (LongMemEval, LoCoMo, STALE, PersonaMem) put 47.4-66.1% of their errors in the retrieval-failure bucket versus only 5.8-13.7% in the post-retrieval-reasoning bucket.
Generative UI
Interfaces generated on the fly by models: dynamic layouts, artifacts, and agent-driven frontends.
tagged in 1 of 97 digest issues, most recently 2026-07-03
Human in the loop
Where people sit in agent workflows: approval gates, escalation, and oversight interfaces.
tagged in 0 of 97 digest issues
Information retrieval
Finding the right information at the right time: ranking, hybrid search, and retrieval over scientific and code corpora.
tagged in 44 of 97 digest issues, most recently 2026-07-22 · 3 papers · 1 explorer section
Digest issues
- The model is the constant; the harness is the variable — 2026-07-22
- A rewritten agent loop beat maximum reasoning effort at 40% of the cost — 2026-07-20
- 96 million tokens "saved" and a 7.6% higher bill — 2026-07-20
- The Bridge Documents Your Reranker Buries — 2026-07-19
- Coding-Agent Gains Are Moving From the Model to the Harness — 2026-07-18
- Harness evolution doesn't beat plain test-time scaling on Terminal-Bench 2.1 — 2026-07-16
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- The agent scaffold is now the contested layer — 2026-06-26
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Enhancing Developer Productivity with Google Colab CLI and Agentic Observability — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
Papers
- Engram: A Bi-Temporal Memory Engine Where a Lean Retrieved Context Beats the Full History — The read path retrieves through four channels in parallel — dense semantic, BM25 lexical, graph n-hop from query entities, and recency/salience — fuses them with Reciprocal Rank Fusion, then applies an 'as-of' filter and an abstention gate; the assembled context is hybrid (conflict-resolved facts plus raw session chunks) because facts alone lose recall.
- Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory — At read time an agentic procedure lets the LLM iteratively call memory tools — inspect intermediate results, expand local context around matches, assemble evidence — rather than take a single top-k step; the agentic variant beats a hybrid summary+BM25 reader (79.3% vs 76.0% on LongMemEval_S).
- T-Mem: Memory That Anticipates, Not Archives — T-Mem retrieves through a top-down topic -> scene -> item cascade, scoring each layer with reciprocal rank fusion over both a shared BM25 lexical index and a per-type dense index (bge-m3). On top of these node indices, four write-time 'trigger' families surface host nodes on the query's behalf: items expose three independently-encoded views (concept-only, bridge-only, joint) and the trigger score is the nan-aware max across views, attributed back to the host item. Scenes/items reached through associative triggers bypass the topic prefilter, because gating them by surviving topics would re-impose the similarity-only neighbourhood the system is built to escape.
Memory consolidation
Distilling episodic traces into durable knowledge and skills, and deciding what an agent should forget.
19 papers · 2 explorer sections · tagged in 0 of 97 digest issues
Papers
- Lifelong Learning of Large Language Model based Agents: A Roadmap — Roadmap framing continual/incremental learning for LLM agents, organizing memory, forgetting, and knowledge-accumulation challenges.
- A-MEM: Agentic Memory for LLM Agents — Zettelkasten-style self-organizing agent memory with dynamic linking and note evolution as new memories arrive.
- Neuron-level Balance between Stability and Plasticity in Deep Reinforcement Learning — Per-neuron stability/plasticity balancing to address the stability-plasticity dilemma in deep RL agents.
- SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment — Benchmark moving beyond episodic task scoring to evaluate self-evolving agents across time, targeting 'episodic amnesia' and static-toolset limits.
- Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment — H2-EMV: LM-based relevance estimation drives selective forgetting under learned natural-language rules updated by user feedback; measures QA accuracy retention against memory-size and query-compute savings.
- Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory — Encodes time as continuous phase rotation (not a discrete label) in temporal KG / agentic memory so validity intervals and obsolescence are representable.
- Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations — Evicts old content past the context window but indexes it with keyword bookmarks so the model can recover evicted memory on demand.
- Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments — Liquid-Glass-Crystal three-phase consolidation governed by an Itô SDE / Fokker-Planck (closed-form Beta stationary dist.); proves convergence + memory-capacity bounds and empirically reports forward transfer and forgetting reductions.
- SCM: Sleep-Consolidated Memory with Algorithmic Forgetting for Large Language Models — Memory architecture drawing on neuroscientific sleep-consolidation principles, pairing offline consolidation with an explicit algorithmic-forgetting operator to bound store growth.
- Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents — Typed semantic memory using information-theoretic retrieval scoring to govern what is retained/surfaced across sessions for long-horizon agents.
- Evolve: A Persistent Knowledge Lifecycle for Small Language Models — Teacher-compiled persistent knowledge store refined through sleep consolidation and usage-driven retention/decay for small local LMs.
- ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems — Replaces system-engineering memory metaphors (paging/flat stores) with a 7-layer neuroscience-grounded architecture incorporating consolidation and forgetting layers.
- STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? — Benchmark of 400 expert-validated conflict scenarios / 1,200 queries (contexts to 150K tokens) probing belief revision over time via three dimensions: State Resolution, Premise Resistance, Implicit Policy Adaptation; introduces 'Implicit Conflict' failure mode and CUPMem write-time-revision baseline.
- EvolveMem: Self-Evolving Memory Architecture via AutoResearch for LLM Agents — Self-evolving memory that treats retrieval infrastructure as mutable, auto-researching its own organization across multi-session operation.
- NeuSymMS: A Hybrid Neuro-Symbolic Memory System for Persistent, Self-Curating LLM Agents — Neuro-symbolic, self-curating memory where symbolic structure governs update/curation of persistent cross-session user knowledge.
- Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents — Frames retention/eviction as a constrained multi-step stochastic optimization over a hard storage budget, with the per-step reward explicitly charging miss penalties, reacquisition delay, and stale-information use; the learned policy keeps a smaller, evidence-denser set rather than greedily filling the budget (LoCoMo budget 128: F1 0.302 at 0.76 occupancy vs Mixed-Score 0.069 at 0.99).
- TokenPilot: Cache-Efficient Context Management for LLM Agents — Lifecycle-Aware Eviction tracks each context segment through three states active -> completed -> evictable, gated on evidence that a sub-task achieved its objective and on residual utility. A completed segment is not immediately purged; it retains its physical cache slots as long as residual relevance to ongoing interactions is non-zero, and only an evictable segment (utility decayed to zero) is removed in a single-pass structural purge. State estimation runs via a lightweight zero-shot validator over a compressed historical view in conservative batches of B turns rather than every step.
- Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge — Across four marker-free evolving benchmarks (code mutation, config migration, dependency bumps, API evolution), plain RAG serves the superseded value 15-40% of the time when forced to answer; MemStrata drives this to ~0% via deterministic supersession rather than any decay or scoring-based forgetting heuristic.
- Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents — Frames consolidation on the complementary-learning-systems analogy (fast episodic plus slow consolidating stores) but states it is motivation only, not a biological model. Consolidation is measured by write economy and bounded drift, not task accuracy: EVAF reaches goal persistence and post-unload recovery of 0.812–0.904 with only 2–3 parametric writes per 200 events (L2 drift ~21–29), while Naive-LoRA writes every event (200 writes, L2 drift ~67 on TinyLlama, ~119 on GPT-2) and still fails the goal layer — so writing everything is not enough. Stability-plasticity is handled with replay plus an L2/EWC-style anchor. The unresolved boundary is stale-memory obsolescence: on public Memora event streams EVAF improves forgetting-absence only 91/222 to 95/222 (McNemar p=0.57, not significant), and negative-gradient/anti-training forgetting variants were unstable — append-only selective consolidation does not solve delete/update validity.
Model Context Protocol
The open protocol connecting models to tools and data sources, and the server/client ecosystem built on it.
tagged in 5 of 97 digest issues, most recently 2026-07-06
Model releases
New model launches and availability changes across frontier and open providers.
tagged in 31 of 97 digest issues, most recently 2026-07-22
Digest issues
- Kimi K3 clears the open-weight bar, and an OpenAI cyber model clears the sandbox — 2026-07-22
- Kimi K3 arrives at 2.8T parameters, and everyone starts routing by price tier — 2026-07-20
- Fable 5 Stays, and the Benchmarks Get a Harder Look — 2026-07-18
- Kimi K3 Prices Like Sonnet, and Everyone Wants the Agent's Terminal — 2026-07-17
- Thinking Machines ships Inkling: 975B parameters, Apache 2.0, and 25K tokens per task — 2026-07-16
- Grok Build CLI uploaded whole repos to a Google bucket; Codex hit 7M users — 2026-07-14
- OpenAI claims a math proof, walks back its launch, and the model race turns to cost — 2026-07-12
- GPT-5.6 lands, and the coding-model price floor drops again — 2026-07-10
- Grok 4.5 Ships as Cursor's First General-Purpose Model, and OpenAI Retracts SWE-Bench Pro — 2026-07-09
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- x402 payments reach both edge networks, and GPT-5.6 Sol Ultra heads to Codex — 2026-07-06
- Sonnet 5 Lands, Fable 5 Returns, and ZCode Undercuts Everyone on Price — 2026-07-06
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Fable 5 completes the task as a test and refuses it as production work — 2026-07-04
- Claude Sonnet 5 lands everywhere at once — 2026-07-01
- The frontier model becomes the supervisor — 2026-06-30
- Open weights became the default coding model the week the frontier got restricted — 2026-06-29
- GPT-5.6 and Mythos 5 ship behind a government access gate — 2026-06-28
- OpenAI ships GPT-5.6 behind a government gate, the same day Washington un-blocks Mythos 5 — 2026-06-27
- Cyber models ship faster than the rules for them — 2026-06-23
- GLM-5.2 takes the open-model crown while agents settle into production — 2026-06-22
- Anthropic resets every usage limit while negotiating Fable back from a US ban — 2026-06-20
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- Two labs put their models on the lab bench — 2026-06-18
- GLM-5.2 cracks the open-weight coding frontier, and Cursor goes to SpaceX — 2026-06-17
- Anthropic shipped the best coding model measured, then the government pulled it — 2026-06-15
- The model layer becomes a regulated surface — 2026-06-14
- The day the US government switched off Fable 5 — 2026-06-13
- OpenAI buys Ona while Fable 5 starts downgrading itself — 2026-06-12
- Fable 5 scores 91, real code scores 13 — 2026-06-11
- Claude Fable 5 arrives at twice Opus pricing, and Cognition's day-old FrontierCode crowns it #1 — 2026-06-10
Multi-agent orchestration
Coordinating multiple agents on shared work: topologies, delegation patterns, shared memory, and production reliability.
tagged in 47 of 97 digest issues, most recently 2026-07-22
Digest issues
- The model is the constant; the harness is the variable — 2026-07-22
- A rewritten agent loop beat maximum reasoning effort at 40% of the cost — 2026-07-20
- A Jacobian Conjecture counterexample, and METR's worst cheating rate yet — 2026-07-20
- 96 million tokens "saved" and a 7.6% higher bill — 2026-07-20
- The Bridge Documents Your Reranker Buries — 2026-07-19
- Coding-Agent Gains Are Moving From the Model to the Harness — 2026-07-18
- Harness evolution doesn't beat plain test-time scaling on Terminal-Bench 2.1 — 2026-07-16
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Claude Tag lands 65% of Anthropic's internal PRs while the World's Fair argues over the outer loop — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- The agent scaffold is now the contested layer — 2026-06-26
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- When the loop closes, verification becomes the job — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
Open models
Open-weight model releases and the ecosystem around running, fine-tuning, and evaluating them.
tagged in 20 of 97 digest issues, most recently 2026-07-20
Digest issues
- Kimi K3 arrives at 2.8T parameters, and everyone starts routing by price tier — 2026-07-20
- Bun-in-Rust ships silently, Gemini slips loudly — 2026-07-19
- Thinking Machines ships Inkling: 975B parameters, Apache 2.0, and 25K tokens per task — 2026-07-16
- Frontier agents ace one task and stall across a stream of them — 2026-07-13
- GPT-5.6 lands, and the coding-model price floor drops again — 2026-07-10
- The frontier model becomes the supervisor — 2026-06-30
- Open weights became the default coding model the week the frontier got restricted — 2026-06-29
- The floor is rising: open models, the coding-speed paradox, and the bill for the buildout — 2026-06-29
- GPT-5.6 and Mythos 5 ship behind a government access gate — 2026-06-28
- Codex usage at OpenAI jumped 56x, and the agent stack started grading itself — 2026-06-26
- Open weights match Opus at half the cost, and ship free in Devin — 2026-06-25
- Anthropic Puts Claude in Slack — 2026-06-24
- Cyber models ship faster than the rules for them — 2026-06-23
- GLM 5.2 edges past Sonnet on 1,000 coding tasks as Claude's error rates spike — 2026-06-22
- GLM-5.2 takes the open-model crown while agents settle into production — 2026-06-22
- Anthropic resets every usage limit while negotiating Fable back from a US ban — 2026-06-20
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- GLM-5.2 cracks the open-weight coding frontier, and Cursor goes to SpaceX — 2026-06-17
- The day the US government switched off Fable 5 — 2026-06-13
- OpenAI buys Ona while Fable 5 starts downgrading itself — 2026-06-12
Reliability
Failure modes, recovery, and durable state for agent systems running unattended or at production scale.
tagged in 4 of 97 digest issues, most recently 2026-07-11 · 15 papers · 1 explorer section
Papers
- ReAct: Synergizing Reasoning and Acting in Language Models — Interleaves chain-of-thought reasoning with tool actions so a model can plan, query external sources, and self-correct — reducing hallucination on decision tasks.
- Is Multi-Agent Debate (MAD) the Silver Bullet? Empirical Analysis in Code Summarization & Translation — Structured multi-agent debate yields minimal-to-inconsistent gains over a strong single-agent baseline on software-engineering tasks.
- Why Do Multi-Agent LLM Systems Fail? (MAST failure taxonomy) — 14 failure modes in 3 categories — specification issues, inter-agent misalignment, task verification — built with an LLM-as-judge pipeline at Cohen's κ=0.88.
- λ_A: A Typed Lambda Calculus for LLM Agent Composition — Well-formedness / termination guarantees for agent composition via a typed calculus.
- TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples — Uses TLA+ counterexamples to repair coordination protocols.
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems — RootSE distills 93 real repository-level agent failures (over 4,500 execution steps) into a diagnosis benchmark whose target is the earliest decisive error step - the single early mistake, such as a misread requirement or flawed plan, whose cumulative consequence is the eventual system failure.
- SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? — A 5-bucket failure taxonomy over 526 agent-attributable failures: implementation failure (41.6%) and timeout (31.4%) dominate, then reward hacking (15.4%), premature termination (7.6%), and poor self-verification (4.0%); long context degrades behavior actively — pass rate falls monotonically with consecutive-duplicate run length (claude-code 41.9%→3.2%) and compaction/summarizer trials pass at 0% vs 8.9% without.
- XFlow: An Executable Protocol Programming System for Reliable Multi-Agent Workflows — The paper makes the failure taxonomy explicit before scaling fan-out: a single agent's hallucinated/malformed/misinterpreted output becomes shared state and corrupts downstream decisions. Concrete failure modes observed: tau3-bench baseline reaches the right end state via an invalid path (treats a parameter clarification as authorization to mutate); CorpusQA fails on the interpretation rule not the retrieved value; SWE-bench baseline submits a patch after local validation already failed. The cloud-edge fan-out is gated: edge workers see only assigned chunks, write only declared outputs, and must pass schema + coverage checks before entering global state.
- AgentArmor: A Framework, Evaluation, & Mitigation of Coding Agent Failures — Decomposes non-adversarial coding-agent failure into three sequential points — forming the correct target (underspecification), pursuing it (capability error), and executing it through the harness (stochastic sampling, context decay) — with a chain-rule risk P(unsafe)=1-(1-f1)(1-f2)(1-f3) and scenarios that isolate each stage, cross-cut by four active modes (greenfield, editing, deployment, monitoring) over 8 scenarios, 20 environments, and 59 transcript templates at n>=500 across three frontier models.
- RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents — Names a failure taxonomy for the lab-to-production gap — fragile fixes (patches pass tests but leave latent bugs), token waste (trial-and-error instead of a planned approach), false confidence (never abstaining on impossible/ambiguous tasks), and broken intermediates (codebase left broken between steps). Baseline ReAct agents fail badly against it: no baseline abstained on any of the 6 impossible tasks, and even disciplined agents abstained correctly only 62% of the time. Recovery is the hardest mode and the one scaffolding does NOT fix — smallest gain of all five pillars, token-waste cut only 34%, doom loops persist when root cause isn't in the error message; the authors conclude recovery may need architectural changes beyond configuration-level frameworks.
- NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems — NOVA's verification cascade checks architecture semantics before expensive training or deployment, catching runnable-but-structurally-invalid candidates that pass unit tests but violate recommender-specific invariants like sequence masking; failed-verification diagnostics become reusable 'forbidden directions' that shape later search.
- Glite ARF: Verifier-Driven Research with Parallel LLM Coding Agents — Glite ARF enforces task-level worktree isolation and immutability with a corrections overlay after an incident where one agent step recomputed and corrupted 20,304 historical training rows across 38 feature sets; completed task folders can't be modified, so fixes are separate, auditable downstream tasks.
- Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software — Across 930,000+ agent-authored pull requests, multilevel models show about half the variance in integration friction survives after accounting for the contribution, its author, size, and agent, and it is a repository-level property. Agent-authored contributions concentrate this repository-level friction roughly twice as much as human ones (ICC 0.30 vs 0.16).
- AgentAbstain: Do LLM Agents Know When Not to Act? — The paper gives an agent-native failure taxonomy of 8 abstention scenarios, organized by when a trigger becomes observable (pre-execution reasoning vs. runtime discovery) and where it resides (query, environment state, or tools). Across 17 frontier LLMs in 4 harnesses the best agent, Gemini 3.1 Pro, reaches only 59.5% Paired Accuracy and 13 of 17 stay below 50%, and abstention capability tracks largely independently of general task-solving skill. The signature failure is post-hoc abstention: an agent commits an irreversible action, such as cancelling a flight before checking rebooking availability, then verbally acknowledges the problem, leaving unrecoverable side effects.
- Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs Single LLMs
Retrieval-augmented generation
Grounding model output in retrieved evidence: dense retrieval foundations, agentic search loops, and RAG system design.
tagged in 0 of 97 digest issues
Scientific search
Discovery over scholarly literature: citation graphs, semantic search, and agentic research assistants.
tagged in 0 of 97 digest issues
Synthetic data
Model-generated training and evaluation data: generation pipelines, quality control, and contamination risk.
8 papers · 1 explorer section · tagged in 0 of 97 digest issues
Papers
- Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization — Surveys the persona concept across role-playing and personalization, distinguishing how personas are constructed, conditioned, and evaluated in LLMs.
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models — Proposes evaluating synthetic-data generators by the Quality–Diversity–Complexity (QDC) makeup of their output, finding quality drives in-distribution generalization, diversity drives OOD generalization, and complexity helps both — with explicit quality–diversity trade-offs.
- Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces — Introduces OmniBehavior, the first user-simulation benchmark built entirely from real-world traces, and uses it to show LLM simulators converge to a 'positive average person' (hyper-activity, persona homogenization, Utopian bias), losing individual differences and long-tail behaviors.
- UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents — A unified tool-learning pipeline that builds a 22k+ tool pool and 390k+ instances by combining 10 public datasets with structurally controlled synthetic trajectories (single/multi-hop, single/multi-turn, serial/parallel), adds an Anchor Linkage mechanism for cross-turn dependencies, and a QAOA evaluation representation.
- Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs — Generates clinically grounded synthetic multi-turn counseling dialogues from structured client psychological graphs, encoding the underlying clinical reasoning rather than just masking real utterances, to sidestep confidentiality constraints on real data.
- EngramaBench: Evaluating Long-Term Conversational Memory with Structured Graph Retrieval — A synthetic long-term memory benchmark of 5 personas, 100 multi-session conversations, and 150 queries spanning factual recall, cross-space integration, temporal reasoning, adversarial abstention, and emergent synthesis, holding the answering model fixed (GPT-4o) to isolate memory architecture.
- A Survey on LLM-based Conversational User Simulation — A dedicated survey organizing the design space of LLM-based conversational user simulators (persona conditioning, goal/intent modeling, behavioral realism, evaluation of simulators themselves).
- VeriSim: A Configurable Framework for Evaluating Medical AI Under Realistic Patient Noise — A configurable simulation framework that injects realistic patient 'noise' (incomplete, inconsistent, distracting user behavior) into evaluation, exposing that strong static-benchmark scores collapse under realistic interaction.
Test-time compute
Spending inference-time computation to improve results: reasoning-intensive retrieval, reranking, and search at query time.
5 papers · 1 explorer section · tagged in 0 of 97 digest issues
Verification
Checking that an agent's output actually satisfies the task: test oracles, property checks, and validation of generated artifacts.
tagged in 2 of 97 digest issues, most recently 2026-06-27