Most AI dashboards measure inputs.
Tokens. PRs opened. Prompts issued. Suggestions accepted. Seats activated.
None of them measure whether any of it worked.
I'm not writing this as someone who solved that. I'm writing it as someone whose organization is in the middle of figuring out what to measure, and who found last week's Reuters reporting on Meta useful precisely because it named the problem we're working through.
What Reuters found
On August 26, Reuters published a special report on Project OT — Organization Transformation — built from scores of internal Meta documents, posts and recordings, plus conversations with more than 20 people with knowledge of the company's operations. (Link)
The plan came out of Zuckerberg's January leadership retreat in Hawaii: an AI-native Meta, where agents absorbed much of the daily work and smaller, talent-dense groups of humans supervised them. Executives had spent the prior year studying how AI-first startups in Asia organized themselves.
The org design is worth reading carefully, because it's more specific than most coverage suggested.
An early pilot run by product VP Ime Archibong set up five small tech pods, each with two to three engineers and a designer, all equipped with AI tools. They dropped fixed six-month planning cycles in favor of four-week prototype sprints. An internal post the following October, titled the AI-Native Playbook, proposed extending this: traditional product designer and engineer roles would disappear, pod members would carry a single generic title — builder — and layers of middle management would be removed, with pods reporting to a single high-level unit head.
The management model that came with it is the part I keep returning to. Unit heads, called Org Leads, would each oversee 30 to 50 people and own performance ratings and promotions. Pod Leads would steer pods day to day but hold no formal management authority. One staffer assigned to lead a pod posted internally that they weren't going through manager training and weren't getting access to ratings or manager tools.
By June, at least 11 units, including engineering and research teams, had implemented small pods.
Then the internal data came in.
Code changes to the internal software platforms and infrastructure employees use on the job: up 220% year over year, from an early-June post by CTO Andrew Bosworth.
Changes that led to new or upgraded features reaching users: up 36%.
Major technical and security incidents, including service disruptions and possible data leaks: up 40%.
Time staffers spent firefighting them: up 70%.
Internal employee sentiment, per Meta's half-year Pulse survey: down from 74% favorable to 55%.
Infrastructure teams had flagged reliability warning signs as early as March. An April post said unchecked AI agents were performing "large-scale, disruptive actions that humans are unlikely to execute." The public got a glimpse in early June, when attackers exploited Meta's new AI-powered customer support bot to reach high-profile Instagram accounts, including the dormant Obama White House page.
Meta declined to comment on the internal data about the AI disruptions.
On the night of May 19, hours before the first layoff wave, Zuckerberg pulled back. Meta cut roughly 10% of its workforce the next morning and cancelled the November wave. In early July he appeared at an internal town hall and conceded the timing had been miscalculated — the agent technology had not accelerated as fast as he'd anticipated, though he expected benefits within three to six months.
One thing to be precise about
Meta did not stop investing in AI. They stopped restructuring around it.
The same reporting notes Meta plans to invest at least $130 billion in AI chips and infrastructure this year, an amount analysts expect will consume its 2026 operating cash. Zuckerberg simultaneously launched a public campaign positioning Meta as people-centric.
That distinction is the whole story. They kept the tool and abandoned the org redesign built on assumptions about the tool.
That's what a functioning second column does. It doesn't tell you to stop using AI. It tells you which bets the current evidence supports.
The part most of the coverage skipped
The story got read as a failure narrative. Big company overreaches, retreats, embarrassing.
That's not the useful reading.
Meta could see the second column. That's why they could stop.
Most organizations cannot make that decision, because nothing in their measurement stack would ever trigger it. If your only instrumentation is adoption — prompts, PRs, seat activation — every quarter looks like progress. The line goes up. There is no signal anywhere in your dashboard that would tell you to pause, because you aren't measuring the thing that goes wrong.
The instrumentation is what bought them the option to change their mind.
The obvious objection
The 220% and the 36% measure different stages of the pipeline. One counts changes to internal platforms and infrastructure. The other counts what reached users. They are not a clean productivity ratio, and anyone presenting them as one is overreaching.
Take that seriously. Then notice what it implies.
Most of us instrument the first number well and the second barely at all. We can report prompts issued and PRs opened to three significant figures. Far fewer of us can say what cleared review, shipped, and then survived ninety days without generating an incident.
The asymmetry is the finding. Not the ratio.
This isn't only a Meta problem
One company, one moment, one set of leaked internal posts. If that were the only evidence, I'd hold it loosely.
It isn't. GitClear has been running a longitudinal analysis of code quality across hundreds of millions of changed lines. Their 2025 report found 2024 was the first year on record where within-commit copy/paste exceeded moved — that is, refactored — code, with refactored lines falling from about 25% of changed lines in 2021 to under 10% in 2024. Their 2026 maintainability research tracks block duplication at the highest level on record.
One finding from that work is worth sitting with. Heavy AI users out-produce non-users by 4–10x — but most of that gap predates AI. Measured against their own past selves, heavy AI users showed roughly a 25% velocity gain.
The 4–10x is the number that ends up in a deck. The 25% is the one that's actually about the tool.
Same input/output confusion, quantified. A metric that looks like it's measuring your AI investment is often measuring who adopted it first.
Training is not the lever we think it is
A study posted to arXiv last week sharpens the people side. Researchers analyzed 713,564 prompts from nearly 4,000 back-office employees across 15 functional areas over eight months in 2025.
Three findings:
Senior employees used GenAI more sophisticatedly, consistent with domain expertise complementing the tool rather than substituting for it.
Sophistication varied sharply by function, highest in Strategy, Digital Innovation, and Project Management.
It did not improve with time on the platform, and formal AI training produced no lasting gains.
That population is back-office knowledge work, not engineering, so don't over-map it. But the direction should give pause to anyone running broad AI literacy training alongside structural downsizing.
If sophistication tracks expertise rather than instruction, then training everyone while shrinking teams is a strange pairing. You'd be diluting the input that drives sophisticated use while scaling the tool that depends on it.
Which connects back to Meta's org design. Pod Leads without manager tooling, Org Leads at 30–50 reports, specialists converted to shared resources — that structure removes exactly the people whose domain expertise makes AI output survivable.
What we're actually working through
We're building the second column now. I don't have results to report. I have a problem statement and a set of arguments.
The candidate signals aren't exotic. Most are things engineering organizations already collect and don't connect:
Change failure rate and time to restore, segmented by whether the change came through an AI-assisted path.
Short-term churn — what fraction of merged code gets substantially rewritten within two weeks.
Review burden per shipped unit. If AI doubles authored volume and review time rises with it, you've moved the bottleneck, not removed it.
Ninety-day survival. What shipped, and what was still standing a quarter later without a rollback or an incident attached.
Incident attribution. Which incidents trace back to changes in the assisted path, and whether that share is growing.
The hard part isn't picking metrics. It's attribution.
You cannot cleanly separate AI-assisted work from the rest. Self-reported tool usage is unreliable. Tool telemetry tells you a suggestion was accepted, not that it survived. And the moment any of this is tied to performance evaluation, the measurement is corrupted — people optimize what's counted, and adoption metrics are trivially easy to game.
That's where we are. If you've solved attribution, I'd genuinely like to hear how.
The gap everyone is misreading
McKinsey's 2025 State of AI survey found 88% of organizations regularly using AI in at least one business function, up from 78% a year earlier — with roughly one-third having begun to scale enterprise-wide.
That gap usually gets read as a technology readiness problem.
I think it's mostly a measurement readiness problem. Most companies can't yet tell whether their AI activity is producing proportionate output or proportionate cleanup. Which means scaling is a commitment made without a control panel.
Meta had the control panel. It cost them a partial reorganization to build it, and when the readings came back, they changed the plan.
The pod model may still turn out to be right. One company, one moment, before those pods had time to accumulate institutional knowledge — that settles nothing about the model.
But it settles a smaller and more immediately useful question. Adoption metrics are not throughput metrics, and an organization that can only see the first column is flying on a broken instrument.
If you're running AI-native teams, squads, or pods: do you have both columns, or only the first?
If you're still in pilot: before you redesign the org, what would it take to instrument what you have well enough to know what you'd be giving up?
We’re in the middle of this one. Comments open.
Sources: Reuters special report on Meta Project OT, August 26, 2026. GitClear, AI Copilot Code Quality (2025) and The Maintainability Gap (2026). Hallman et al., "Sophistication in GenAI Use: Field Evidence from a Large Firm," arXiv:2608.27364. McKinsey, The State of AI: Global Survey 2025.
