khalil

📍 paris

blog

A place where I share my unfiltered thoughts

On building a useful personal harness

tech ai engineering

TLDR: you should have a cross-project harness, compounded on project specific harness appendices.

There's been a trend of people open sourcing their harnesses, gstack ↗ by YC's Garry Tan being the famous example. But they almost all miss the mark on the one attribute that is key to really building a useful harness: specificity.

Building a personal harness is inherently an iterative process, a constant self reflection loop where one must ask « where am I, the human, slowing things down ». Which pieces of the workflow can be factored into skills, loops or graphs? And if this process is not married with a specific problem, then it ends up being too vague and unapplicable.

In my experimentation I found it much more useful to separate my harnesses in two categories.

A cross-project harness, which centralizes all the skills, tools, scripts and roles that are general enough to be used across multiple projects. Typically feature spec & definition with roles (product manager, tech lead, designer, qa), skills tailored to project management in Linear (roadmaps, what to pick up next), and GitHub workflows (e.g. Claude ↔ Codex review loops).

An appendix harness for each project, which builds on the cross-project one and extends scope and spec contracts with project specific goals and design rules. In practice this is one small config file per repo. It says which document defines the project scope, and what new skills or tools to add to the project.

cross-project harness defined once, pointed at by every repo roles pm · tech-lead · designer · qa review loops claude ↔ codex · land planning linear · roadmaps · next general skills, no project nouns project A .agents/stack.yml scope doc · design rules skills · tools project B .agents/stack.yml scope doc · design rules skills · tools project C .agents/stack.yml scope doc · design rules skills · tools
One harness, an appendix per project. The specificity lives in the appendix, not in the skill.

The cross-project layer is organized by phase, not by tool. A request gets routed to exactly one skill, and each skill sits in the phase where it is useful.

request decide /spec · /triage /next build /dispatch /investigate review /pr-loop claude ↔ codex land /land stops at the PR operate /health · /retro /linear-steward ↑ the human merges what the retro finds becomes next cycle's scope
The five phases of the cross-project layer. The only place a human is mandatory is the merge.

Now, running multiple sessions in parallel can get a bit out of hand, so here are 3 rules I force myself to follow to keep track of all of the project context:

  • At any given time there's a scope document, or an ADR, I can refer to when deciding if a new feature gets built. If nothing in the roadmap asks for it, the answer is no by default. If the new feature helps the roadmap but is not explicitly planned, I invoke the “spec” skill to evaluate the tradeoffs in including it.
  • The human has the final say on the merge. Agents open the PR, they never merge it. This is what keeps things from getting out of hand. Even when I'm not deep in the review, I still have to look at the evals and the unit tests, and see real evidence.
  • Merging PRs and closing tickets is not enough. The only thing that matters is whether the eval, scoped against the roadmap, passes or not.

When used appropriately with a clear scope (a detailed 2 weeks roadmap), this setup truly feels like operating at the speed of a small startup team.

The harness is open source: kstack ↗

A Jungian perspective on Sensui

misc psychology anime