Agentic Engineering: What It Is, and the Acceptance Problem It Creates

October 1, 2026 · Kealu Vector Team · Engineering

Agentic engineering defined, and why production capacity and acceptance capacity run on different clocks in software whose failure has consequences.

Agentic engineering is the practice of having software agents perform engineering work under human direction: planning, implementing, verifying and documenting changes for which a named person remains accountable.

An agent plans a change, writes it, runs it, reads the failure, revises. Somebody then decides whether the result can be accepted, and answers for that decision afterward.

Producing the work has become fast. Deciding to accept it moves at the speed of reviewers, evidence and named authority. When the second lags the first, the difference accumulates as work that exists and cannot ship.

What is agentic engineering

That definition is the one this article uses throughout, and the term is already in circulation with a close meaning. Andrej Karpathy, who coined "vibe coding" in early 2025, drew the distinction at Sequoia Ascent in April 2026. In the

A note on that source, since this article is about accepting generated work. The same page carries a summary produced by a language model from the video, and the sentences most often quoted from this talk come from that summary rather than from the transcript. The two quotations above are from the transcript section, which Karpathy describes as cleaned up for transcription errors and filler. The distinction cost one reading of the page to establish, and it is the same class of check this article argues for.

The definition above keeps that discipline and adds the part a regulated organization cannot skip: the acceptance is somebody&039;s, by name, and it has to be evidenced.

Three parts of the definition carry weight.

Nothing in the definition depends on how capable the agent is. A more capable agent produces more work per hour. The question of who accepts that work, and on what basis, stays where it was.

Two capacities, running on different clocks

A working repository is an instrument. The Kealu white paper

The paper states the shift in one line: "AI is shifting the economic bottleneck in software engineering from production to acceptance." It names the two capacities an engineering organization runs on.

> Production capacity refers to the ability to generate designs, code, tests, documentation, analyses, and changes.

> Acceptance capacity is the ability to determine whether a particular outcome is suitable for a stated use, supported by relevant evidence, and accepted by an accountable authority.

The paper is explicit that "Although related, these capacities are not interchangeable." Production capacity scales with the agent population. Add another agent, give it a task, and output arrives. Acceptance capacity scales with the number of people qualified to judge a class of change, the evidence those people can get without asking for it, and the speed at which an accountable decision can be made and recorded. Those inputs grow slowly. A reviewer who can judge a timing change in a braking system takes years to develop, and an evidence trail that a second engineer can read without reconstructing the work has to be produced while the work happens.

The paper puts the consequence in one sentence: "An organization can increase its production capacity and reduce delivery performance if review queues, integration burdens, evidence gaps, or unresolved decisions grow faster than the output."

A large-scale measurement is consistent with that attenuation. Demirer, Musolff and Yang matched development records to AI-usage telemetry for more than 500,000 GitHub developers and estimated the effect of each generation of coding tool in a matched event study.

!

The figure is a hypothesis. Its assumptions, written out:

A curve with no numbers makes a claim about mechanism, which is the part we would like the field to test. Anyone with measurements that contradict the shape has something we want to read.

The maturity illusion

The gap has a way of staying invisible while it opens. The white paper calls this the maturity illusion.

> Agentic tools can produce visible functionality quickly; therefore, projects appear to advance sharply. However, system maturity depends on integration, non-functional behavior, evidence, operational constraints, and convergence across disciplines. If these obligations are deferred, the gap is discovered later when architectural flexibility is lower and change is more expensive.

Every status report reads well while this happens. Features demo. Tests pass. The deferred obligations are the ones nobody demos: whether the timing assumptions still hold at system level, and whether an auditor two years from now can follow what was decided and why.

Where acceptance goes wrong

Obvious failures are cheap to reject. The white paper locates the expensive output elsewhere: "Difficult output is coherent, plausible, and almost correct."

The 2025 Stack Overflow Developer Survey

The same survey asked how far developers trust the accuracy of AI output. Its summary of the answers:

> More developers actively distrust the accuracy of AI tools (46%) than trust it (33%), and only a fraction (3%) report "highly trusting" the output. Experienced developers are the most cautious, with the lowest "highly trust" rate (2.6%) and the highest "highly distrust" rate (20%), indicating a widespread need for human verification for those in roles with accountability.

The white paper lists the shapes this takes in engineering work.

> It may satisfy the prompt while violating an unstated rule. It may generate tests that confirm the assumptions embedded in its implementation. It may use a valid library in an unsupported manner. It may handle the normal case while missing the timing, concurrency, rollback, or degraded-mode requirements.

The output is good enough that the constraint has moved downstream of it. Any capable contributor working without the unstated context produces the same class of defect. Agents change the volume and the speed at which it arrives.

The paper itemizes the bill: the cost of converting that into accepted engineering work "is paid through senior review, architectural reconstruction, integration debugging, test expansion, compliance mapping, and operational caution." None of it appears at the moment the agent finishes.

The measurement problem

It is tempting to settle all of this with a productivity number. One randomized trial, and its sequel, show how hard that is.

In July 2025, METR published a randomized controlled trial. Sixteen experienced open-source developers worked on 246 real issues from repositories they had contributed to for years, each issue randomly assigned to allow or disallow AI. The result: developers

> This gap between perception and reality is striking: developers expected AI to speed them up by 24%, and even after experiencing the slowdown, they still believed AI had sped them up by 20%.

Self-assessment of how much faster the work went proved unreliable even for experienced engineers in familiar code.

METR now labels that result out of date. In February 2026 the team

> An increased share of developers say they would not want to do 50% of their work without AI, even though our study pays them $50/hour to work on tasks of their own choosing. Our study is thus systematically missing developers who have the most optimistic expectations about AI&039;s value. [...] When surveyed, 30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI.

> Based on conversations with study participants, we believe it is likely that developers are more sped up from AI tools now [...] compared to our estimates from early 2025. However, because of the selection effects in our experiment, our data is only very weak evidence for the size of this increase.

One thing follows for acceptance. A production-speed number, on its own, makes a weak basis for deciding how much agent-produced work an organization should take on, and METR&039;s own account shows how quickly such a number goes stale. The evidence an organization needs sits in its own queue, its own review record and its own defect history.

The five conditions acceptance runs on

A decision about suitability needs material to decide on. The white paper calls for "an independent Assurance Plane that governs how agentic engineering work is performed and evidenced", which "must preserve five conditions: independent verification, validation against authoritative intent, named human authority, transferable evidence, and evidence freshness." That Assurance Plane is the thing Kealu names the Agentic Engineering Assurance Layer. Our reading of each condition:

| Condition | The question it answers | What it produces | | --- | --- | --- | | Independent verification | Who checked this, and how far from the producer do they stand? | A check whose result does not depend on the assumptions of the work it is checking | | Validation against authoritative intent | Was the right problem solved? | A comparison of the result against the authoritative intent and intended use | | Named human authority | Who accepted this, and on what basis? | A person or authorized role attached to the decision, with the uncertainties they were shown | | Transferable evidence | Can somebody else reach the same conclusion later? | A record another qualified person can read without reconstructing the work | | Evidence freshness | Is the evidence still valid today? | A link between each claim and what it depends on, so a change can flag what went stale |

Independence is the condition that is easiest to appear to satisfy. An agent that writes an implementation and then writes the tests for it has verified very little: the tests inherit the misunderstanding. What makes a check worth reading is separation between the thing that produces and the thing that checks, and how much separation a piece of work needs is a risk decision rather than a fixed setting. The white paper gives the range: "Depending on the consequence, independence may require a separate agent, model, method, tool, environment, organization, or human specialist."

The bottom of that range is the part most often mistaken for the whole of it. A different model or a different prompt changes who does the checking without necessarily changing what the checking rests on: the same intent, the same unstated assumptions, the same context assembled the same way. That is separation of execution, not separation of judgment, and it is not by itself a guarantee of independent judgment. At the integrity levels where a failure costs lives, the separation that counts sits at the far end of the range and is organizational: an organization other than the one that produced the work. Between those two ends the answer is a risk decision, proportionate to consequence and to the customer&039;s own policy.

Evidence has a shelf life. A dependency update, a standard revision, a change in a neighboring component, and an assurance claim that held in March is unsupported by September, with nothing in the system saying so.

Consequential Software: where the gap gets expensive

For much software, a change that ships in the morning, is observed by lunch and is reverted in a minute carries its own verification. The acceptance question gets expensive in a specific class of system, which the white paper names.

> Consequential Software is software whose failure, degradation, manipulation, unavailability, or untraceable change can materially affect human safety, health, financial integrity, essential services, legal obligations, or the operation of a complex system, and whose acceptance therefore requires lifecycle evidence and an accountable technical authority beyond local functional correctness.

The paper classifies by mechanism: a useful classification "asks how consequences are produced." Five dimensions do most of the work, and a sixth matters for long-lived products.

The paper works the point through three composite scenarios, one of them from financial infrastructure. An agent changes the retry logic of a payment-processing service, adds backoff and idempotency keys, and "runs unit and integration tests, and they pass." Then: "A retry that is safe before authorization may be unsafe after a partially completed settlement step". The operational dashboard can show successful retries while a downstream ledger records duplicate economic events. The risk lives in a relationship outside the visible feedback surface. A qualified human or an independent system has to reconstruct that relationship before anybody accepts the change.

What a team can change without buying anything

Most of the acceptance problem is organizational, and a team can start on it with the tools it already runs.

All six run on tools a team already has, once the record is treated as part of the work.

Questions we have not settled

We put this category forward to be tested.

If you run Consequential Software and have views on any of these, we would like to hear them.

Frequently asked questions

Where this goes next

The white paper

The Agentic Engineering Assurance Layer is the proposed category for the layer that governs how agentic engineering work is performed and evidenced, so that humans can trust and accept the resulting work. The companion white paper develops it as a vendor-neutral category, and the conditions above hold whatever tools a team uses to meet them. Kealu Vector, currently designated Vector Assurance Assist, is our early-access implementation of that category. What this article says about it describes an intended approach, not measured product results. This article is about the problem itself.

Related articles