Cara LiNotes on AI, work & capital

Essay 01 · AI, organizations & investing

From Job Replacement to Organizational Compression

How AI reshapes firms, investment funds, and private equity portfolios

AI agents may not yet be able to replace a single employee end to end, but they may already enable a much smaller team to perform the work of a much larger one. Agents' Last Exam, led by Berkeley RDI, covers more than 1,500 long-horizon tasks across 55 professional fields. OpenAI reports that GPT-5.6 Sol achieved a partial-credit score of 53.6%, while the ALE leaderboard reports an end-to-end pass rate of only 30.6%. In other words, the model can complete meaningful portions of professional work without reliably finishing the entire workflow. In the near term, AI may therefore compress the teams surrounding employees rather than replace entire jobs.

The workflow, not the job, is the better unit of analysis. A job typically combines information collection, analysis, coordination, exception handling, judgment, and accountability, and these activities do not become automatable at the same rate. Agents are already effective at gathering information, conducting first-pass analysis, drafting documents, and updating systems. Humans remain essential for complex judgment, unusual cases, and final accountability. As a result, the traditional preparer-reviewer-approver structure may evolve into an agent-expert model, with people focused on exceptions and responsibility.

A useful way to understand this unevenness is through falsification latency: how long it takes to discover that a decision was wrong. Tasks with fast, objective feedback—such as reconciliations, formatting checks, and unit tests—can support greater agent autonomy. Work whose errors become visible only through expert review—such as contracts, complex analysis, and client deliverables—still requires human approval. In strategy, system architecture, and investing, the relevant verifier may not arrive for months or years. Tests can confirm that an output works today without confirming that the underlying decision will survive future conditions. The longer the falsification latency, the more organizations must preserve human ownership of decomposition, judgment, and accountability.

For companies and investors, the real economic value lies not simply in helping employees perform existing tasks faster, but in redesigning the workflow itself. Handoffs and repetitive reviews can be removed, organizational knowledge can be converted into a reusable knowledge base, and expert feedback can continuously improve agent performance. Technical capability becomes economic value when internal friction and noise dial down.

Organizational compression, however, creates a longer-term challenge: where will the next generation of experts come from? Traditionally, junior employees developed judgment through years of seemingly repetitive hands-on work. In the future, firms may need to expose a smaller group of trainees to a deliberately designed mix of common cases, difficult problems, rare exceptions, and examples of both good and bad decisions. The best compressed organizations will reduce the cost of execution without weakening the development of human judgment.

From Reasoning to Action

Organizational compression is only the first-order effect. The deeper opportunity is to turn execution itself into a learning system. Companies can turn their workflows into training environments by giving agents access to the necessary tools, permissions, verifiers, and feedback. Every completed task then produces proprietary data about what worked, what failed, and where human intervention was required. This creates a continuous loop of execution, evaluation, and improvement.

The next AI moat may therefore lie less in the underlying model than in the quality of the surrounding system: the workflow environment, evaluation criteria, accumulated feedback, and integration between training and deployment. The companies that learn fastest and at the lowest cost from real-world outcomes will improve their agents fastest.

Reinventing the Fund: From Research Factory to Learning System

Investment funds have traditionally been organized around research production: junior analysts collect information and maintain models, senior investors form views, and portfolio managers allocate capital. As agents automate information gathering, model updates, first-pass analysis, and monitoring, the scarce capabilities shift toward asking the right questions, evaluating evidence, identifying variant perception, sizing positions, and bearing responsibility for outcomes.

An AI-native fund is not a traditional fund that produces more research; it is a decision system. That system should learn continuously through a loop of question, evidence, thesis, position, outcome, and attribution. P&L is an important but incomplete signal: it reveals the outcome without showing which parts of the investment process were right or wrong. A profitable investment may reflect weak reasoning and good luck, while a losing investment may reflect a sound thesis with poor timing or sizing. A verifier layer cannot eliminate long feedback cycles, but it can structure them and shorten the time required to recognize error. The fund therefore needs predictions and falsifiers recorded before investment, separate evaluation of thesis accuracy, timing, and sizing, and postmortems that improve future decisions. The next great fund may not produce the most research, but learn fastest which judgments deserve capital.

Reinventing the PE Firm: From Deal Teams to a Portfolio Learning System

Private equity firms have an advantage that public-market investors do not: they can change how portfolio companies operate. An AI-native PE firm can instrument high-volume workflows, define measurable outcomes, deploy agents on bounded tasks, and reserve human judgment for exceptions and accountability. Each portfolio company then becomes both a cash-flow asset and a learning environment.

The resulting loop is: acquire, instrument, redesign workflows, deploy agents, verify outcomes, and replicate what works across the portfolio. Workflow data, evaluation criteria, and operating lessons accumulate at the fund level, allowing each transformation to improve the next one. The moat is not a single cost-reduction program, but the ability to convert operating experience into a reusable knowledge base and transformation playbook.

About the authors

Cara Li portraitCara Li

Cara Li is an MBA student at Harvard Business School. Before HBS, she worked across data, evals, productization, and enterprise deployment on Zhipu AI’s LLM Research team, after starting her career in investment banking at Goldman Sachs. She is interested in how AI reshapes organizations, workflows, and capital allocation.

LinkedIn
David Han portraitDavid Han

David Han is an AI PhD researcher at BAIR, UC Berkeley. He builds benchmarks and environments to measure whether AI agents can perform long-horizon, economically valuable work, and develops methods that enable them to ultimately outperform domain experts.

Personal website
← Back to all writing