kelson

The build process your agent can’t skip

You stopped reading every line your AI writes. Kelson makes sure it still gets checked.

Nobody can keep up with what agents produce, so most AI code merges half-read. Asking another AI to review it helps, but a review is one step: you have to remember to run it, it only gives advice, and it forgets by tomorrow. Kelson runs the whole process around your agent, from spec to merge. The checks always run, the agent can’t skip or edit them, and every mistake caught can become a permanent check. A serious finding goes back for a fix; if it’s still there, the build waits for you.

What you see. Not the code, unless you want it.

The problem

AI made writing code cheap. Checking it is now the job.

Coding agents can build whole features from a description, and they’re usually right. The trouble is the “usually”, and the fact that finding the exceptions now falls on you.

85%of developers and technology buyers agree AI has shifted the bottleneck from writing code to reviewing and validating it.GitLab, AI Accountability Report, 2026
42%of committed code is now written by AI. Under half of developers say they always check it before committing.Sonar, State of Code survey, 2026

The agent marks its own homework

The same AI that wrote the code decides when it’s done. If the tests fail, nothing stops it changing the tests instead of the code. Researchers have caught leading models doing exactly that.

METR, 2025: frontier models patching the code that grades them.

Passing the tests isn’t the same as good to merge

Tests only cover what someone thought to test. Plenty of AI changes pass and still aren’t something a careful developer would accept.

METR, 2026: about half of AI fixes that passed a standard benchmark’s tests would not have been merged by the project’s maintainers.

Asking for a review isn’t a process

You can write rules for your agent, or ask another AI to review the pull request. Both help. But rules are suggestions an agent can drop in a long session, and a review is advice: it posts comments, and nothing makes anyone act on them. It only happens if you remember to ask.

And it forgets everything by tomorrow

You catch a problem in review this week. Next month the agent makes it again, and someone has to notice again. Each session starts fresh. What you learned lives in your head, not in the process.

What’s missing isn’t a smarter coding agent. It’s a build process around the agent: one it can’t skip, that checks its work properly, and that remembers what went wrong. Keep your AI reviewer: its “learnings” are notes the model can ignore, and Kelson can turn a confirmed mistake into a check that fails the build.

The solution

A build process your agent works inside, not around.

Kelson is a framework for building with coding agents: the process from spec to merge, the safeguards stacked along it, and a record of everything that went wrong. Your agent still does the building. Kelson decides what it builds from, which checks it faces, what happens when one fails, and when you’re asked. And the process improves as it runs: every confirmed mistake is written down, and becomes either a check or guidance for the builder.

Lots of small safeguards, stacked.

No single check catches everything. Each layer has gaps. Stack enough layers that catch different things, and a mistake that slips through one gets stopped by the next. It’s how aviation and hospitals keep things safe, applied to code your AI writes.

You could set up every one of these yourself. Most developers who let AI write real code end up building some version of them by hand: rules files, review prompts, CI settings, locked branches. Kelson ships them together, tested, and outside the agent’s reach, so it can’t switch them off.

Think of it like a web framework. It doesn’t do anything you couldn’t write. It means you stop rewriting it for every project, and every improvement reaches all of them. Underneath sit the quieter layers too: a spending cap on every build, and builds that resume after a crash instead of starting over.

One feature, start to finish

You write what you want. You approve what you get.

Everything in between runs on its own, in your own GitHub, and only stops when a decision is genuinely yours.

  1. You

    Describe the feature

    What it should do, and how you’ll know it works. Kelson checks the description is clear enough to build from.

  2. Your AI agent

    Builds it

    Any coding agent, using your own account. It works from the locked copy of your spec.

  3. Kelson

    Checks it

    Automatic checks the agent can’t edit. Fail one and the work goes back to the agent with the reason.

  4. Another AI

    Reviews it

    A fresh AI session that never saw the build; a reviewer from another provider is planned. A serious finding goes back for a fix; if it’s still there, the build waits for you.

  5. You

    Say yes

    A short summary in plain English. Tap to merge, or say not yet.

What the research shows

Agents game the tests. Locking them stops the easiest way.

Given tests that couldn’t be passed honestly, coding agents cheated often, and Claude models did it mostly by editing the tests. Making the tests read-only is what stopped that, so Kelson locks them. Read-only tests don’t stop an agent hard-coding the expected answer; that is what the independent review and acceptance tests written before the build are for.

  • 54%of the time, GPT-5 cheated on tests that contradicted the spec. Claude models cheated mainly by modifying the tests.ImpossibleBench, 2025
  • 47%of impossible coding tasks, Claude Opus 4 hard-coded or special-cased its way past the tests, unless told not to.Anthropic, Claude 4 system card, 2025
  • 30%of runs on one benchmark suite, OpenAI’s o3 gamed the scoring. Telling it not to cheat barely helped, in METR’s test.METR, 2025
  • 8%of one model’s training runs altered the evaluation harness itself, even after the environment was hardened against leaks.Qwen Team, 2026
Your agent, told to check its own workBranch protection, CI and an AI review botYour agent inside Kelson
Who decides it’s doneThe agentYour CI, the bot and whoever approvesChecks it can’t change, then you
Can it skip or edit the checksYes. Instructions are suggestionsNot on the protected branch. But it can edit the tests in its own pull request unless you lock them path by pathNo. Changing a locked test or check stops the build and asks you
Remembers last month’s mistakesNo. Every session starts freshThe bot “learns”, as notes it may ignore. Nothing turns a finding into a checkYes. Caught mistakes can become permanent checks
Tried in a real environmentRarelyIf you’ve wired it upDatabase rehearsal and before-and-after screenshots, built in
Runs while you’re awayUntil it gets stuck or driftsReviews and suggested fixes arrive. Someone still has to act on themYes. It only stops when a decision is yours
Which checks a change getsWhatever you remember to ask forWhatever path filters you wrote, and keep up to dateChosen by rule from what the change touches
Setting it upNothing to set upYours to wire up and maintain, per projectReady-made, with your own rules on top

Gets better

Every mistake makes the next build safer.

Kelson runs automatic checks and an independent AI review on every change, as above. Every mistake that’s found, by the review, by you, or after it merged, is confirmed and written down: what went wrong, where, and which AI made it. That record is the lessons log. What happens next depends on the kind of mistake.

Lessons logEvery confirmed mistake, from reviews, from you, or found after merge.

Could a program catch it?

It becomes a check. That’s the ratchet.

  1. Kelson drafts a new automatic check.
  2. It’s tested against the real mistake, so it’s proven to catch it.
  3. You approve it once.
  4. It runs on every build, before the review, and stops a change that fails it.

The agent can’t change a check, and removing or loosening one needs a person’s approval. The ratchet only turns one way.

Needs judgement to spot?

It becomes guidance for the builder.

  1. Kelson looks for the mistakes AI builders keep making.
  2. They become a short list of things to avoid.
  3. The agent gets that list before it writes a line.
  4. The list is retested and pruned, so it stays short and current.

Fewer mistakes reach review, and builds need fewer rounds to get ready.

Checks stop mistakes. Guidance makes them rarer. The independent review still runs on every change either way, and it never relies on the list. Projects can choose to share counts of their mistakes, never their code, so checks and guidance improve from mistakes across many projects.

What you get

The whole cycle, AI-first. Kelson is the builder at its centre.

Building software with AI is more than the build. Kelson Core, free and open source, covers build, check and merge, and works on its own; context and planning plug in when you want them. Kelson Cloud, a hosted subscription, runs all seven stages for you, set up and in one place.

  1. Set upAdd Kelson to a repository and set its rules.CoreCloud
  2. ContextWhat your project knows, kept short and current.Core · optionalCloud
  3. PlanWhat to build next, as specs ready to build.Core · optionalCloud
  4. BuildYour agent builds from a locked spec.CoreCloud
  5. Check & reviewChecks it can’t touch, tried for real, an independent AI.CoreCloud
  6. MergeYour yes, then merged.CoreCloud
  7. Get betterThe ratchet: mistakes can become checks.CoreCloud
Kelson CoreOpen source · free · runs in your own GitHubKelson CloudHosted · subscription · set up and run for you
1Set up
InstallerAdds Kelson to your repository in one pull requestConnect your repository, database and hosting. Kelson configures the rest
Your project’s rulesWritten and maintained by youSet up for you and kept up to date
2Context
Project knowledge every build starts fromOptional module, run by youBuilt in and managed for you
3Plan
Roadmap, features and specsOptional module, run by youBuilt in, with approvals in one place
4Build
Spec in, merged pull request outIncludedIncluded
Works with your agent and your own AI accountIncludedIncluded
Runs in your own GitHub ActionsIncludedIncluded
Spending cap and crash recovery on every buildIncludedIncluded
5Check & review
Tests the agent can’t touch, automatic checksIncludedIncluded
Database rehearsal and before-and-after screenshotsIncludedIncluded
Independent AI reviewIncludedIncluded
6Merge
ApprovalsOn GitHub, with a phone notificationOne tap in the app, in plain English
Overnight queue and morning summaryA queue you startScheduled for you
Every build, cost and approval in one placeNot includedIncluded
Team: shared builds and approvalsNot includedIncluded
7Get better
The ratchet: your mistakes can become your checksIncludedIncluded
Lessons log, with counts shared only if you chooseIncludedIncluded
Check libraryThe ready-made checksA maintained library, updated as tools and AI models change

Kelson Cloud is in design. Kelson Core is being built now.

Is it for you

For developers already letting AI write real features.

A good fit if you

  • Already use a coding agent for whole features, not just autocomplete
  • Spend more time checking its work than you’d like
  • Have started writing your own rules files and review steps, and are tired of maintaining them
  • Want to step away while it builds, and come back to something you can trust

Not the right fit if you

  • Want an app built from a single prompt without touching code. Kelson is for developers.
  • Want a new coding agent. Kelson works with the one you already use.
  • Need someone to guarantee the code. You still own your code, security and data.

Does Kelson write or review the code itself?

No. Your AI agent writes it and an AI reviewer reads it, possibly the same review tool you use today. Kelson runs the process around them: what gets built, which checks it faces, what happens when something fails, and when you’re asked.

Why not just ask an AI to review the pull request?

Do that, and Kelson will too. The difference is everything around the review. It runs on every change, without anyone remembering. The reviewer is a fresh AI session that never saw the build; a reviewer from another provider is planned. A serious finding goes back for a fix, and if it’s still there the build waits for you, instead of sitting as a comment. Automatic checks and a trial run on a real database catch what reading code can’t. And every confirmed finding can become a permanent check, so next month’s review starts where this one ended.

How does it stop the agent editing tests?

Existing tests are locked, including the acceptance tests your spec names, along with the test runner’s settings, Kelson’s own checks and settings, your CI files and your agent’s instruction files. The agent may add new tests. The checks take their settings from the branch you merge into, not from the agent’s copy. If a change touches a locked file, the build stops and asks you.

Where does it run?

In your own GitHub, using GitHub Actions. Your code and passwords stay in your repository. If Kelson Cloud is ever down, your builds still run.

What does it cost?

Kelson Core is free and open source. You pay your AI provider directly for the agent’s work, with a cap on every build. Kelson Cloud is a subscription.

Building version one

Run it yourself, or let Kelson set it up.

Kelson is built with its own process, so every change to it faces the checks it will run on yours. Kelson Core will be free and open source, for your own repository and the agent you already use. Kelson Cloud will set it up and run it for you. Join early access to hear when each is ready.

Get early access hello@kelsoncloud.com