AI agents: harnesses, context, memory and dreaming
A model's behavior depends on the software around it, the information it receives, and what it retains between tasks. These notes compare several ways to improve that system, with links to the original articles, specification and technical report.
Sources checked on 9 October 2026. Experimental numbers below are author-reported results, not independent reproductions.
Harnesses and context engineering
An agent harness runs the interaction loop around a model. It provides tools, constructs context, handles tool results and decides when to retry or stop. Context engineering determines what information reaches each model call, including instructions, retrieved documents, tool output and selected memories. Persistent memory stores information for later tasks; retrieval decides which parts belong in the current context.
The ultimate guide to multi-harness RL, by Adithya S Kolavi and colleagues, describes training inside existing harnesses with OpenEnv, Harbor and TRL. A capture proxy records generated tokens and their probabilities for reinforcement learning.
The authors trained LFM2.5-2.6B on 1,000 tasks and evaluated 250 held-out tasks under OpenCode, Claude Code, Codex and Mini-SWE-Agent. Average pass@1 rose from about 42% to 54.2%, with 31% fewer tool calls on tasks already solved by the base model. OpenCode-only training reached 52.3%; the authors say the 1.9-point overall gap is within noise. Transfer across individual harnesses differed. These task results do not establish general improvement across all coding benchmarks.
Clement Delangue's announcement describes the same training approach. Its full long-post text and outbound article link remain unverified. The independently retrieved guide is a web research article; its PDF export requires Pro access, and the authors identify the web version as canonical.
Learn harness engineering with WalkingLabs
An agent writes plausible code, announces it is finished, and leaves the next session to discover the missing tests. WalkingLabs' Learn Harness Engineering course starts with this gap between model capability and reliable execution.
“The agent says ‘I'm done’ when it's not”
The course turns that problem into practical work: state clear completion criteria, give agents the right tools and project conventions, preserve progress across sessions, and verify the result. Lectures lead into hands-on projects and reusable templates for instructions, feature lists and progress notes.
Start with Project 01: prompt-only versus a minimal harness. Run a task with and without explicit rules and verification, then compare what actually works. This is an exercise you can reproduce in your own repository, rather than a promise that a particular harness will always improve performance.
Devin memory and Agent Memory Repo
Learning across sessions
Cognition's Memory and dreaming: how Devin learns from working with you, dated 5 October 2026, describes a personal Git repository of Markdown memories. A short MEMORY.md supplies general preferences and an index; sessions retrieve other notes as needed. Separate checkouts and revision checks protect concurrent updates.
Daily background dreaming reviews conversations and existing notes, merges duplicates, removes transient or stale material, and discovers lessons across sessions. Source references connect memories to earlier work. Skills package repeatable procedures; memory records lessons accumulated while using them. This is a product description without a controlled performance benchmark. The Cognition announcement links to the proposed format.
A proposed open memory format
Agent Memory Repo specifies an entry-point file, notes with optional source and date metadata, and wiki-style links between files. Agents clone the repository, retrieve relevant information, edit notes and push updates. Periodic dreaming can consolidate records and resolve contradictions.
Cognition describes it as an open standard. Here it is treated as Cognition's proposed open memory format, without assuming standards-body endorsement or broad adoption. The specification and implementation repository is open to contributions. Its local trial excludes automatic startup and scheduled dreaming. The website's agent-swarm and SQL examples illustrate intended use; they are not experimental evidence.
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Tong Zheng and colleagues, with affiliations at Google, the University of Maryland, Google DeepMind and the University of Virginia, describe Dream-RSI in a project page, technical report PDF and code repository.
Dream-RSI holds the discovery agent fixed and improves executable exploration policies. Online work produces discovery trees with recorded execution outcomes. Offline replay tests alternative search decisions against those trees; the selected policy returns online and expands the history. Dreaming here changes exploration code, whereas Devin dreaming curates text memory.
The project reports 2.43 times fewer generations on VGG16 at comparable performance, and 2.09 times higher performance on ConvDiv at comparable budgets. Its headline 162-fold reduction in Lasso discovery-agent calls compares 317 calls with SimpleTES's 51,200, using different models. That comparison does not isolate the exploration policy's contribution.
Replay covers only branches already explored. Selecting the best candidate, including the incumbent, prevents a lower selection score on recorded history; it does not guarantee improvement on unseen tasks. The site's statement that an agent must dream to self-improve is the authors' framing, not a demonstrated universal requirement.
Wake-sleep for legal agents
Niko Grupen's post links to The Return of Wake-Sleep. Public post metadata verifies that title, link and opening preview. The following details come from the supplied Pasted text.txt, whose complete article body was not independently retrieved.
The supplied account describes an agent working on legal tasks, then an offline review model deriving lessons from graded traces. It merges overlapping lessons, rejects client-specific or overly narrow entries through a Jev commit gate, and stores reusable checklists and practice notes for later tasks. This updates textual memory rather than model weights.
The authors report ten cycles across 196 Legal Agent Benchmark tasks in Corporate M&A and Capital Markets, with 110 training tasks and 86 validation tasks, using GPT-6 Luna as the agent. The all-pass rate rose from 2.9% to 15.7%, a 12.8 percentage-point increase. They report rubric-level gains on both familiar and new matters.
Memory also prompted about 2.5 times more analysis, checking and drafting tool calls, increasing cost and latency. Retrieving relevant lessons instead of supplying the full memory reportedly halved per-task cost while preserving the rubric-criteria pass rate. That efficiency comparison is against full memory, not the no-memory baseline.
These are rubric-judged benchmark results from the supplied account. They do not establish reliability on real client matters. Confidence intervals, independent replication and the full experimental artifacts remain unverified here.
Evidence limits and X threads
The sources improve different parts of an agent system. Multi-harness RL changes model weights; memory dreaming revises stored notes; legal wake-sleep curates task lessons; Dream-RSI revises exploration policies. Their results use different tasks, baselines and cost measures, so the numbers should not be ranked against one another.
Engineering questions follow from those differences. Does a remembered lesson retain its source and scope? Can contradictory or stale notes be corrected? Does retrieval preserve quality while reducing cost? Does an improved policy transfer beyond the history used to select it? Answering these requires held-out evaluation and accounting for the offline work as well as online execution.
Highest-liked replies remain unresolved
The requested X threads did not expose a verifiable collection of reply bodies and like counts. No highest-liked reply summary is presented. A defensible ranking needs reply text, engagement counts and a collection timestamp; post popularity alone cannot establish which replies ranked highest.
Cognition's announcement and Niko's article link were verified through public post metadata. Clement's accessible text is truncated. The guide above matches its topic and claims, but the exact outbound link in the complete post remains unknown.
Original sources and further reading
- WalkingLabs: Learn Harness Engineering. A practical course with lectures, hands-on projects and reusable templates covering agent boundaries, continuity across sessions, verification and observability. It is a learning resource, not a controlled evaluation of agent reliability.
- Devin memory and dreaming article and Cognition's announcement.
- Agent Memory Repo format and open repository.
- Clement Delangue's post and multi-harness RL web article. The complete post-to-article link remains unconfirmed.
- FineEnvs article and training resources and TRL's experimental harness-training documentation.
- Dream-RSI project, technical report and code. The site's arXiv identifier is a placeholder; the direct report is the verified paper link.
- Niko Grupen's post and The Return of Wake-Sleep, supported by the supplied text attachment.
Related notes on this site: multi-agent learning path, clinical reasoning and AI, and rfab-harness for antibody design campaigns.