← all repositories
walkinglabs/awesome-harness-engineering

Because even Claude Code needs a seatbelt

A curated index of articles, playbooks, and tools for harness engineering—the discipline of shaping the environment around AI agents so they stop drifting on long-running tasks.

3.7k stars AgentsLLMOps · Eval
awesome-harness-engineering
Velocity · 7d
+15
★ / day
Trend
accelerating
star history

What it does This is an awesome list that catalogs resources defining “harness engineering,” the practice of building scaffolding around AI agents to keep them reliable during long-running workflows. It collects vendor field reports, academic position papers, and open-source tools from OpenAI, Anthropic, Thoughtworks, and others into a single taxonomy covering context management, guardrails, evaluation, and observability. The list explicitly excludes generic agent frameworks unless they address reliability-critical primitives like runtime control or state management.

The interesting bit Rather than treating agent failures as model problems, the list frames them as harness problems—arguing that context windows, sandboxing, and handoff artifacts matter more than raw prompting. It attempts to formalize a scattered set of vendor best practices into a coherent discipline with its own vocabulary, including the “control–agency–runtime” decomposition and structured reporting via HarnessCard.

Key highlights

  • Heavyweight source material from OpenAI, Anthropic, LangChain, and Thoughtworks on context engineering and safe autonomy.
  • Practical patterns for repo-local instructions (CLAUDE.md, AGENTS.md, init.sh) and spec-driven agent workflows.
  • Coverage of eval and observability tooling, including Inspect AI and OpenTelemetry semantic conventions for generative AI.
  • A dedicated section on constraints and guardrails, from sandboxing to mitigating prompt injection in autonomous coding agents.
  • Links to a companion course repository with practical harness projects built around an Electron knowledge-base app.

Caveats

  • As a curated index, it contains no original code or implementations; it is purely a reading list and signpost.
  • The README acknowledges one broken survey link with a <!-- FIXME --> comment, suggesting maintenance is ongoing but not perfect.
  • Several sections list competing or overlapping standards (e.g., AGENTS.md vs. agent.md) without reconciling them.

Verdict Worth bookmarking if you are building or operating long-running coding agents and want to move beyond vibe-coding into reproducible, observable systems. Skip it if you are looking for a drop-in framework or a quickstart tutorial.

Frequently asked

What is walkinglabs/awesome-harness-engineering?
A curated index of articles, playbooks, and tools for harness engineering—the discipline of shaping the environment around AI agents so they stop drifting on long-running tasks.
Is awesome-harness-engineering open source?
Yes — walkinglabs/awesome-harness-engineering is an open-source project tracked on heatdrop.
How popular is awesome-harness-engineering?
walkinglabs/awesome-harness-engineering has 3.7k stars on GitHub and is currently accelerating.
Where can I find awesome-harness-engineering?
walkinglabs/awesome-harness-engineering is on GitHub at https://github.com/walkinglabs/awesome-harness-engineering.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.