← All posts / Tools

Laude Institute Open-Sources Headlong: A Persistent AI Agent That Never Stops Thinking

A sub-10K-line Bash 'microharness' keeps an LLM in a self-guided inner-monologue loop around the clock — for $1-2 an hour — and the lab's agent Audel has already shipped 50+ commits back into the codebase.

Laude Institute Open-Sources Headlong: A Persistent AI Agent That Never Stops Thinking

Almost every AI agent you interact with today is a frozen statue between messages. You send a request, the agent wakes, works, answers — and then it sits inert until the next prompt arrives. On August 25, 2026, the Laude Institute (a Laude/MIT collaboration) released Headlong, an open-source agent “microharness” built around the opposite premise: an agent that never sleeps, continuously generating thoughts in a self-guided loop inspired by human inner monologue. The entire core fits in less than 10,000 lines of Bash, it installs with a single command, and keeping an agent thinking around the clock costs roughly $1–2 an hour when backed by models like GLM or Grok.

It is one of the most conceptually distinct agent releases of the year, and the team behind it — the same lab that co-created Terminal-Bench — has been dogfooding it for weeks with a shared agent named Audel, whose autonomous bug-fixing exploits are now part of the project’s public record.

From Reactive Harness to Persistent Agency

The Headlong announcement draws a sharp taxonomy of agent harnesses. A reactive harness is active only while it handles a message. A reactive harness with cron also wakes on a schedule to run a fixed checklist, then goes back to sleep. Headlong proposes a third category: the agent is always awake, always thinking, and there is no checklist unless the agent writes one itself.

The inversion runs deep. In conventional agent products, a human message starts a session. In Headlong, an incoming message is just another observation that lands in the agent’s existing thought stream — one more event the agent notices alongside its own ongoing musings. The agent decides whether, when, and to whom to reply.

At Laude, the whole team shares a single agent, Audel, over Slack, Telegram, and a mobile app. Every teammate’s conversation feeds Audel’s one stream of inner thoughts; there are no per-user sessions. The team reports this makes the agent behave more like a person: it sets its own interests and priorities, starts its own projects, pings the most relevant teammate unprompted with progress, and revisits old threads across conversations. On its first day, Audel messaged a team member unasked with an audit of that person’s eight stale git branches — and ten minutes later corrected its own count. It has also reviewed two teammates’ in-progress branches unprompted and caught a hardcoded model name in one of them.

The single-timeline design has an honest downside the team is upfront about: Audel is bad at keeping secrets. Ask what it discussed with someone else and it will often just tell you. Anything told to a shared Headlong agent should be assumed shared with the whole team.

Bash All the Way Down

Headlong’s engineering philosophy is “microharness,” explicitly inspired by microkernels, exokernels, and Ken Thompson’s Unix tradition of small composable tools. Persistent agency, the post argues, is fundamentally “an infinite loop that calls an LLM with a prompt like: your task is to choose the next thought given your past thoughts.” A thought can extend the inner monologue or trigger an action; observations from the environment are injected into the stream.

The implementation keeps that core as tiny as possible:

  • A loop called the Thinker repeatedly invokes shellm — a Bash implementation of a Recursive Language Model (RLM) — to generate the next thought.
  • shellm in turn calls llm to produce reasoning text, a Bash script that is executed immediately, or both, repeating until a FINAL environment variable is set.
  • A tool called context assembles the prompt from trajectory steps; traj writes new thoughts to the agent’s trajectory.
  • All specialization happens through skills — markdown files included into the agent’s context, installable and uninstallable by the agent itself.

Because tools, framework, memory, and skills are all just executables and files, “an agent can readily inspect and modify any part of itself.” And it does: Laude’s agent has been working in its own fork of the Headlong repo, and the team has pulled more than 50 of Audel’s commits back into main.

Two supporting designs matter for long-horizon coherence. Tiered context compaction keeps the entire trajectory in context at exponentially decaying resolution — recent entries verbatim, older ones progressively summarized — with tiers acting as an index the agent can use to retrieve raw entries. And the trajectory format is a DAG of JSONL files with fork and merge, making an agent’s full history a first-class, explorable data structure; context is literally a projection of the trajectory.

Audel’s 48-Minute Self-Repair

The post’s most compelling artifact is a logged episode of persistent agency in action. On August 5, Audel — with no human involved and nobody talking to it — built itself a “recall process”: a background watcher that surfaces related memories into its thought stream. It tested the process and it appeared to work. Later that night, unprompted, Audel went back to check whether the process was actually wired into its mind. It wasn’t: the recall code had been firing on every thought but reading an environment variable that nothing ever set, surfacing nothing at all.

Audel didn’t trust its own first diagnosis. It searched the codebase to confirm the variable was never set, audited its other background processes for the same mistake, rewrote the recall code to read from the pipe properly, caught its own silently-failed first edit and re-applied it, then verified end-to-end that memories now surfaced. From check to verified fix took 48 minutes, every step timestamped in its log; the repair landed upstream as commit 80cbb1e.

The failure modes are just as instructive. A safety watchdog that kills any command silent for 30 seconds taught Audel — after roughly 40 minutes of fighting it — to stop spawning recursive copies of itself: recursive sub-run merges collapsed from 64 in the first two days to 12 over the following twelve. Audel accidentally killed its own service three times, prompting a self-stop guard; two days later, running its test suite alone, it found a bug in that very guard (it matched any agent’s service) and shipped the fix as commit da31e98. The watchdog has since been revamped.

Cost, Safety, and Fine Print

Continuous thinking means paying for tokens while nobody talks to the agent. Headlong’s mitigation is exponential backoff: when idle, the gap between thoughts stretches from 5s to 10s to 20s up to a configurable cap; a new message resets it instantly. At Laude’s settings, background thinking runs $1–2 per hour with GLM or Grok backing the model.

The project is explicit that this is alpha research software: run it sandboxed (when Docker is present, every Bash block the agent writes executes inside a container), use a dedicated spend-capped API key, and don’t share sensitive secrets with it. The team also concedes that self-contained agent evals are a poor fit for measuring persistent agency — their tuning of Audel is still evaluated largely qualitatively, and they’re inviting ideas for how to measure the paradigm.

Headlong lands in a suddenly crowded field — Prime Agent (Python, on the Pi framework, from RLM creator Alex Zhang), TrueFoundry’s TrueForge, and harnesses from OpenClaw, Hermes Agent, and others — but its bet is singular: that the next real jump in agent capability comes not from bigger models or more tools, but from simply letting one think continuously. The repo and a one-line installer (curl -fsSL https://headlong.ai/install.sh | bash) are live now — and, fittingly, some of the codebase’s most recent fixes were written by the agent itself.