October 1, 2026 · Kealu Vector Team · Engineering
Agentic engineering defined, and why production capacity and acceptance capacity run on different clocks in software whose failure has consequences.
Agentic engineering is the practice of having software agents perform engineering work under human direction: planning, implementing, verifying and documenting changes for which a named person remains accountable.
An agent plans a change, writes it, runs it, reads the failure, revises. Somebody then decides whether the result can be accepted, and answers for that decision afterward.
Producing the work has become fast. Deciding to accept it moves at the speed of reviewers, evidence and named authority. When the second lags the first, the difference accumulates as work that exists and cannot ship.
That definition is the one this article uses throughout, and the term is already in circulation with a close meaning. Andrej Karpathy, who coined "vibe coding" in early 2025, drew the distinction at Sequoia Ascent in April 2026. In the A note on that source, since this article is about accepting generated work. The same page carries a summary produced by a language model from the video, and the sentences most often quoted from this talk come from that summary rather than from the transcript. The two quotations above are from the transcript section, which Karpathy describes as cleaned up for transcription errors and filler. The distinction cost one reading of the page to establish, and it is the same class of check this article argues for. The definition above keeps that discipline and adds the part a regulated organization cannot skip: the acceptance is somebody&039;s, by name, and it has to be evidenced. Three parts of the definition carry weight. Nothing in the definition depends on how capable the agent is. A more capable agent produces more work per hour. The question of who accepts that work, and on what basis, stays where it was. A working repository is an instrument. The Kealu white paper The paper states the shift in one line: "AI is shifting the economic bottleneck in software engineering from production to acceptance." It names the two capacities an engineering organization runs on. > Production capacity refers to the ability to generate designs, code, tests, documentation, analyses, and changes. > Acceptance capacity is the ability to determine whether a particular outcome is suitable for a stated use, supported by relevant evidence, and accepted by an accountable authority. The paper is explicit that "Although related, these capacities are not interchangeable." Production capacity scales with the agent population. Add another agent, give it a task, and output arrives. Acceptance capacity scales with the number of people qualified to judge a class of change, the evidence those people can get without asking for it, and the speed at which an accountable decision can be made and recorded. Those inputs grow slowly. A reviewer who can judge a timing change in a braking system takes years to develop, and an evidence trail that a second engineer can read without reconstructing the work has to be produced while the work happens. The paper puts the consequence in one sentence: "An organization can increase its production capacity and reduce delivery performance if review queues, integration burdens, evidence gaps, or unresolved decisions grow faster than the output." A large-scale measurement is consistent with that attenuation. Demirer, Musolff and Yang matched development records to AI-usage telemetry for more than 500,000 GitHub developers and estimated the effect of each generation of coding tool in a matched event study. ! The figure is a hypothesis. Its assumptions, written out: A curve with no numbers makes a claim about mechanism, which is the part we would like the field to test. Anyone with measurements that contradict the shape has something we want to read. The gap has a way of staying invisible while it opens. The white paper calls this the maturity illusion. > Agentic tools can produce visible functionality quickly; therefore, projects appear to advance sharply. However, system maturity depends on integration, non-functional behavior, evidence, operational constraints, and convergence across disciplines. If these obligations are deferred, the gap is discovered later when architectural flexibility is lower and change is more expensive. Every status report reads well while this happens. Features demo. Tests pass. The deferred obligations are the ones nobody demos: whether the timing assumptions still hold at system level, and whether an auditor two years from now can follow what was decided and why. Obvious failures are cheap to reject. The white paper locates the expensive output elsewhere: "Difficult output is coherent, plausible, and almost correct." The 2025 Stack Overflow Developer Survey The same survey asked how far developers trust the accuracy of AI output. Its summary of the answers: > More developers actively distrust the accuracy of AI tools (46%) than trust it (33%), and only a fraction (3%) report "highly trusting" the output. Experienced developers are the most cautious, with the lowest "highly trust" rate (2.6%) and the highest "highly distrust" rate (20%), indicating a widespread need for human verification for those in roles with accountability. The white paper lists the shapes this takes in engineering work. > It may satisfy the prompt while violating an unstated rule. It may generate tests that confirm the assumptions embedded in its implementation. It may use a valid library in an unsupported manner. It may handle the normal case while missing the timing, concurrency, rollback, or degraded-mode requirements. The output is good enough that the constraint has moved downstream of it. Any capable contributor working without the unstated context produces the same class of defect. Agents change the volume and the speed at which it arrives. The paper itemizes the bill: the cost of converting that into accepted engineering work "is paid through senior review, architectural reconstruction, integration debugging, test expansion, compliance mapping, and operational caution." None of it appears at the moment the agent finishes. It is tempting to settle all of this with a productivity number. One randomized trial, and its sequel, show how hard that is. In July 2025, METR published a randomized controlled trial. Sixteen experienced open-source developers worked on 246 real issues from repositories they had contributed to for years, each issue randomly assigned to allow or disallow AI. The result: developers > This gap between perception and reality is striking: developers expected AI to speed them up by 24%, and even after experiencing the slowdown, they still believed AI had sped them up by 20%. Self-assessment of how much faster the work went proved unreliable even for experienced engineers in familiar code.Two capacities, running on different clocks
The maturity illusion
Where acceptance goes wrong
The measurement problem