← All articles
AI Agents Are Quietly Becoming the New Org Chart

GrowthJun 28, 2026

AI Agents Are Quietly Becoming the New Org Chart

The real story isn't autonomous AI doing your job — it's companies using agents to discover how decisions actually get made, then redesigning themselves around the answers.


TL;DR

  • AI agents are moving from coding assistants to organisational mirrors: the more interesting deployment isn't just writing code or answering tickets, but using agents to surface the implicit rules that govern how a company actually operates [5][6].
  • The tooling layer is maturing fast: Claude Code, Codex, Braintrust's eval-and-CI pipeline, and Replit's internal agents are all shipping real work — not demo decks — in production environments [1][2][7][8].
  • The strategic frame is the differentiator: HBR argues that companies must translate tacit principles into structured guidance for agents, and that agent deployment will function as an X-ray of organisational design [5][6].
  • This is a Tier-2/Tier-3 story with real novelty but thin corroboration: the ideas are fresh and largely non-overlapping with prior coverage, but no Tier-1 primary source anchors the claims, and most are single-source.

What happened

A cluster of recent pieces — two from Harvard Business Review, several from Lenny's Newsletter, and one from SaaStr — collectively describe a shift in how businesses are thinking about AI agents. The shift is from tools that do tasks to systems that encode organisational decision-making.

On the tooling side, Lenny Rachitsky's "How I AI" series covers writing AI agent loops in Claude Code and Codex, including an episode where an agent identified a 15-year-old bug in Mozilla Firefox [1]. A companion piece reviews Claude Fable 5 and explores how Braintrust uses AI agents, evals, and continuous integration to ship better software [2]. Claire Vo's interview with Ankur Goyal of Braintrust elaborates on the same theme, with Goyal arguing that "evals are the modern version of a PRD" — a claim that positions evaluation frameworks as the new spec document for AI-mediated work [7]. Fiona Fung, who manages the Claude Code and Cowork teams at Anthropic, discusses what happens after coding is "solved" — a question that implicitly reframes the engineering function itself [4].

On the strategic side, HBR publishes two pieces that move the conversation beyond tooling. K. Sudhir argues that companies need to design agentic systems around the implicit rules governing their company — the unwritten norms, heuristics, and trade-offs that actually drive decisions [5]. Jen Stave extends this: companies must teach their AI how they make decisions, translating tacit principles into structured guidance that agents can follow [6]. The punchline, stated plainly in the source material: "The firms that win will use agent deployment as an X-ray and redesign their organizations around what they find" [5].

SaaStr's piece adds a concrete deployment example: Jason Lemkin and Amjad Masad (Replit's co-founder and CEO) presented live at SaaStr AI 2026, showing "the actual AI agents doing the actual work" — including an AI VP of Marketing called "10K" and an AI customer success rep called "QBee" — explicitly framed as "not a demo deck" [8].

What it actually means

The throughline across these sources is that AI agents are becoming a forcing function for organisational clarity. This is the real story, and it's more interesting than the surface narrative of "AI does your job now."

Consider the coding-agent coverage. The detail that an agent found a 15-year-old bug in Firefox [1] is a neat hook, but the structural point is that agents like Claude Code and Codex are being used to tackle deeply technical architecture and infrastructure work that no single human engineer could tackle before [1]. This isn't automation of routine tickets; it's expansion of the problem-space a single team can address. Braintrust's approach — pairing agents with evals and CI [2][7] — formalises the feedback loop: agents generate work, evals encode taste and correctness, and CI enforces the standard. Goyal's framing of evals as the modern PRD [7] is telling: it suggests that the specification is moving from prose written for humans to test suites written for machines. When the spec becomes executable, the role of the engineer shifts from writing code to defining and guarding the spec.

Fiona Fung's question — what happens after coding is solved? [4] — pushes this further. If the implementation layer becomes cheap, the scarce resource is judgement: deciding what to build, why, and against which trade-offs. This connects directly to the HBR pieces. Sudhir's argument is that every company runs on implicit rules — the senior leader who always overrides the roadmap in Q4, the unwritten rule that customer complaints from enterprise accounts get escalated within two hours, the heuristic that a feature ships when the founder says it does [5]. These rules are invisible in the org chart and absent from most process documentation, but they are the actual operating system of the company. Agents, to be useful, need to be taught these rules — and the act of teaching them forces the company to articulate what it actually believes [6].

This is why the X-ray metaphor matters. Deploying an agent isn't just a technology decision; it's a diagnostic. The agent will fail in ways that reveal where your decision-making is incoherent, contradictory, or entirely dependent on a single person's tacit knowledge. The firms that win will use that failure as a map of what to fix [5]. The firms that lose will blame the agent.

The SaaStr example [8] is the most concrete instantiation of this. Lemkin and Masad showed agents with names and roles — 10K as AIVP of Marketing, QBee as AI Customer Success — running actual SaaStr operations. The framing of "not a demo deck" is doing heavy lifting here: it's a claim that these are production agents, not investor theatre. Whether they're genuinely autonomous or heavily supervised is an open question the source doesn't fully resolve, but the intent is clear — the org chart now includes non-human roles.

Hype deconstruction

This story is not "AGI is here and your job is gone." None of these sources claim general autonomy. The agents described are bounded: they operate within loops, are constrained by evals, and require structured guidance to function [1][2][5][6][7]. The Firefox bug-find is impressive but is a single instance, not a benchmark [1]. The SaaStr agents have names and titles, but the source provides no performance data, no error rates, no comparison to human baselines [8].

This is also not a story with strong evidentiary foundations. The signal score reflects thin corroboration: only one claim in the bundle is corroborated across multiple sources (the general framing of AI agent deployment in business), and the rest are single-source [1][2][3][4][5][6][7][8][9][10]. There are no Tier-1 primary sources — no peer-reviewed benchmarks, no independent audits, no regulatory filings. The HBR pieces are original reporting but are essentially framework essays, not empirical studies [5][6]. The Lenny's Newsletter pieces are practitioner interviews and community roundups [1][2][3][4][7][9][10]. SaaStr is a conference write-up [8].

So the hype to resist is twofold. First, don't over-index on the autonomy framing. These agents are tools that require specification, evaluation, and oversight — the more interesting story is how that specification changes the organisation, not how the agent replaces the worker. Second, don't treat practitioner narratives as evidence of generality. Braintrust's eval-and-CI pipeline works for Braintrust [2][7]; SaaStr's named agents work for SaaStr [8]. These are n=1 deployments in companies with strong engineering cultures and high AI fluency. The leap to "this works everywhere" is exactly the leap the sources do not support.

Stakeholder landscape

Engineering leaders (CTOs, VPs of Eng, staff engineers) are the primary audience for the coding-agent coverage [1][2][4][7]. The Vo/Goyal interview is explicitly pitched at senior engineers who need to demystify evals and use them to improve AI products without touching implementation [7]. Fung's piece targets the same group but asks the harder downstream question: if coding is solved, what is the engineering function for? [4].

Founders and operators are the audience for the HBR pieces and the SaaStr write-up. Sudhir and Stave are writing for leaders who must make strategic decisions about agent deployment [5][6]. Lemkin's SaaStr piece is founder-to-founder: here's what we actually built, here's what Replit's founder thinks comes next [8].

Product managers appear implicitly throughout. The claim that evals are the modern PRD [7] directly affects how PMs define scope. The Community Wisdom pieces [3][9] touch on product operating models and team-level context-sharing, suggesting the PM function is already being reshaped by agent deployment.

AI vendors — Anthropic (Claude Code, Cowork), Replit, Braintrust — benefit from this coverage because it normalises their tools in production settings. The Lenny's Newsletter network functions as a distribution channel for these vendors' narratives: each "How I AI" episode is effectively a case study for a specific product [1][2][4][7].

Workers whose tacit knowledge is being encoded are the least visible stakeholders but the most affected. If an agent can be taught how a senior leader makes decisions [6], the leader's unique judgement becomes less unique. The X-ray metaphor cuts both ways: it reveals organisational incoherence, but it also commodifies the knowledge needed to fix it.

Cross-layer implications

One non-obvious connection: the eval-as-PRD framing [7] and the HBR argument about implicit rules [5][6] are the same idea at different scales. At the product level, evals encode taste and correctness for a specific feature. At the organisational level, structured guidance for agents encodes taste and correctness for how the company decides. Both are moves from tacit knowledge to executable specification. This suggests a convergence: the same engineering discipline — writing tests that capture intent — is becoming the discipline of management. The implication is that the companies best at CI/CD may also become the companies best at organisational design, because they already have the muscle for encoding judgement into systems.

A second connection: the SaaStr agents with names and roles [8] point toward a future where the org chart is a mixed human-AI document. This raises governance questions none of the sources address. Who is accountable when 10K (the AIVP of Marketing) makes a bad call? What does performance management look like for QBee? The sources describe deployment but not governance — and that gap is where the next wave of friction will come from.

What this means for you

If you're a senior engineer or engineering leader, the actionable signal is concrete: start treating evals as a first-class artefact, not an afterthought. Goyal's claim that evals are the modern PRD [7] means your evaluation suite is becoming the contract between your team and your agents. Invest in it the way you'd invest in a spec doc. The Braintrust model — agents, evals, CI [2][7] — is a reference architecture worth studying.

If you're a founder or operator, the HBR framing is your action item. Before deploying an agent, try to write down the implicit rules it would need to follow [5][6]. Where you can't articulate them — where you find yourself saying "it depends" or "we just know" — that's your X-ray. Fix the incoherence before you deploy the agent, or the agent will expose it for you.

If you're a product manager, the eval-as-PRD shift [7] means your role is moving from writing prose specs to defining testable intent. This is a skill transition, not a job elimination — but it's a real one.

If you're an individual contributor whose tacit knowledge is the thing being encoded, the strategic question is whether you're the person doing the encoding or the person being encoded. The former is a position of leverage; the latter is a position of risk.

Uncertainty ledger

  • No Tier-1 primary sources: the entire bundle is Tier-2 and Tier-3. The HBR pieces are frameworks, not studies [5][6]. The practitioner pieces are interviews and conference write-ups [1][2][4][7][8]. Independent benchmarks or audits would materially change confidence.
  • Single-source claims dominate: nearly every specific claim — the Firefox bug, the SaaStr agents, the eval-as-PRD framing, the X-ray metaphor — rests on one source [1][5][6][7][8]. Corroboration would strengthen the analysis significantly.
  • No performance data: the SaaStr agents are described as doing real work [8], but no metrics are given. How autonomous are they? What's the error rate? How do they compare to human baselines? Unknown.
  • Generalisability is untested: Braintrust's pipeline and SaaStr's agents are n=1 deployments in high-AI-fluency organisations. Whether the patterns transfer to companies with weaker engineering cultures is an open empirical question.
  • Governance gap: none of the sources address accountability, liability, or performance management for agents operating in production roles [8]. This is a known-unknown that will become acute as deployment scales.

Bottom line

The real story isn't that AI agents are taking over — it's that deploying them forces companies to articulate how they actually work, and the act of articulation is itself the transformation. The firms that win will use agent deployment as an X-ray and redesign their organizations around what they find [5]. The firms that lose will deploy agents, get bad results, and blame the technology instead of the org chart.

Sources

  1. Lenny Rachitsky. (22 June 2026). 🎙️ How I AI: How to write AI agent loops in Claude Code and Codex + How Claude Mythos found a 15-year-old bug in Mozilla Firefox. lennysnewsletter.com.
  2. Lenny Rachitsky. (15 June 2026). 🎙️ How I AI: Claude Fable 5 review & How Braintrust uses AI agents, evals, and CI to ship better software. lennysnewsletter.com.
  3. Kiyani. (20 June 2026). 🧠 Community Wisdom: Fractional CPO compensation, free e-signature tools, why some users pay but never use your product, sharing Claude Code context across a team, and more. lennysnewsletter.com.
  4. Lenny Rachitsky. (21 June 2026). What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams). lennysnewsletter.com.
  5. K. Sudhir. (19 June 2026). How to Design Agentic Systems Around the Implicit Rules that Govern Your Company. hbr.org.
  6. Jen Stave. (25 June 2026). Teach Your AI How You Make Decisions. hbr.org.
  7. Claire Vo. (15 June 2026). How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal. lennysnewsletter.com.
  8. Jason Lemkin. (25 June 2026). Amjad Masad and Me at SaaStr AI 2026: The Agents We Actually Built, and What Replit's Founder Thinks Comes Next. saastr.com.
  9. Kiyani. (13 June 2026). 🧠 Community Wisdom: How AI is changing product operating models, tracking work stress with Whoop, whether you need a portfolio of AI side projects, marketing for tiny teams, and more. lennysnewsletter.com.
  10. Joe Hudson. (23 June 2026). The new inner game: Your unfair advantage in the age of AI. lennysnewsletter.com.