HomeBook › Future Development
Chapter 12

Future Development

From GenAI for Business (2026 Second Edition) by Shubin Yu · Open in the interactive reader · Download the full PDF

Learning Objectives

After this chapter, you should be able to:

The frontier of generative AI moves faster than any single chapter can capture, so this final chapter looks ahead. We start with what I consider the most consequential shift on the horizon: the rise of World Models. Where Large Language Models mastered language, World Models aim to master reality itself, giving machines an internal simulator of the physical world. Business leaders should resist filing this under distant research curiosities. It is the foundation for physical intelligence, embodied agents, and enterprise-scale digital twins.

12.1 The World Model Frontier: Redefining Intelligence from Pixels to Presence

AI research is turning away from the "lexical mirroring" of Large Language Models (LLMs) and toward physical intelligence, and World Models are the bridge. They are, in a real sense, the prerequisite for the next era of Artificial General Intelligence (AGI). Predictive text mimics the output of intelligence. World Models internalize the mechanics that produce it: physics, spatial relationships, causal dynamics. An AI with that grounding stops being a passive container of knowledge and becomes something that can navigate and manipulate the physical world.

12.1.1 The Genesis of the World Model: From Mental Maps to AI Simulations

The Conceptual Foundation. The genesis of this field lies in "Mental Model" theory, posited by Kenneth Craik in 1943. Craik suggested that human cognition relies on small-scale internal simulators that allow us to anticipate consequences before acting. We do not need to drop every glass to know it will shatter; our internal physics engine rehearses the event in latent space. This mental rehearsal is what makes planning possible. The brain evaluates outcomes without paying the price of physical failure.

Formal Definition. The modern technical architecture was formalized in the 2018 framework by Ha and Schmidhuber, which breaks the World Model into three interlocking modules. The Observation module (V) acts as the system's vision, compressing high-dimensional sensory input, millions of raw pixels, into a manageable latent representation. The Prediction module (M) functions as the memory of the system: a differentiable physics engine that anticipates future world states based on current observations and prospective actions. The Controller (C) is the action module, optimizing policy by "dreaming" inside the simulator created by the Memory module and executing the best solution in the real world only after virtual validation. The conceptual pillars have stood for decades. What changed recently is the arrival of massive video datasets and H100-scale compute, which together push simulation from theory into working practice.

12.1.2 Lexical vs. Physical Intelligence: Why World Models Surpass LLMs

The "Moravec Paradox" remains the primary barrier to entry for trillion-dollar robotics and autonomous markets: high-level reasoning is computationally cheap, but low-level sensorimotor skills are immensely difficult to automate. LLMs, despite their linguistic prowess, are "word smiths in the dark." They possess statistical knowledge of gravity but lack the grounded, physical understanding required to navigate it. World Models change the question itself: not which word is probable, but what motion is caused.

Comparative Analysis: LLMs vs. World Models

Dimension Large Language Models (LLMs) World Models (WMs)
Primary Unit Tokens (discrete semantic units) Pixels / Voxels (continuous visual/spatial units)
Core Goal "Saying things" (Lexical Intelligence) "Seeing and Doing things" (Physical Intelligence)
Data Dependency Static text / image datasets Dynamic / sequential video and sensor data
Understanding Indirect (statistical correlation) Direct (causal and physical grounding)
Hardware Req. High-memory H100 clusters for inference Real-time edge inference for 3DGS / robotics

The "Black Box" Critique. LLMs function as containers of human culture, yet they remain untethered from reality. They can describe a trajectory but cannot simulate it. This lack of spatial intelligence leads to a performance plateau in embodied AI. To reach AGI, systems must move from observing "what follows a word" to predicting "what happens to a voxel" when a force is applied. Without that grounding, an AI stays a stochastic parrot, however fluent it sounds.

12.1.3 The Tripartite Technical Landscape: Three Approaches to World Modeling

The current "technological arms race" is a trade-off between visual fidelity (skin), structural understanding (skeleton), and computational efficiency (nervous system).

Approach A: Video Generation as Simulation (The "Skin" Layer). Models like Google DeepMind's Genie 3 treat pixel prediction as a proxy for world logic, generating interactive 720p environments at 24 FPS with a visual memory that can extend up to 60 seconds. The limitation is that this approach masters the appearance of reality, predicting the next pixel from statistical probability, without necessarily understanding the 3D geometric skeleton underneath. The result looks consistent but carries no explicit spatial coordinates.

Approach B: 3D Native and Spatial Intelligence (The "Skeleton" Layer). Championed by Fei-Fei Li's World Labs (Marble), this approach prioritizes explicit 3D structures, outputting meshes and spatial coordinates using technologies like 3D Gaussian Splatting (3DGS) instead of 2D frames. The advantage: 3DGS is faster and easier to edit than traditional voxels, and when you know exactly where an object sits in a coordinate system, you can plug it straight into physics engines like Unreal or Unity. That gives robotics a much sturdier "skeleton" to work with.

Approach C: Abstract Latent Reasoning (The "Nervous System" Layer). Yann LeCun's JEPA (Joint-Embedding Predictive Architecture) argues that predicting every pixel is computationally wasteful and instead works in latent representations, ignoring environmental noise (the specific glint of light on a leaf) in favor of causal structure (the direction the leaf is falling). For high-level decision-making and planning this is far more efficient, because it strips away redundant data to find what actually drives a scene.

12.1.4 Technical Deep-Dive: NVIDIA Cosmos and the 3D Paradigm Shift

NVIDIA's Cosmos platform marks the move from passive generative AI to a "Digital Twin 2.0" infrastructure. It connects pixel generation to production-ready interactive assets.

The architectural pieces of these World Foundation Models build on one another to turn raw video into a usable representation of reality. Scene perception uses advanced tokenization to convert high-dimensional visual data into compact semantic tokens, allowing the model to "understand" 360-degree panoramas and complex scene layouts with discrete precision. On top of this, trajectory planning applies causal predictive intelligence to navigation: Cosmos imagines multiple future trajectories and picks the path that respects learned physical constraints.

Two further capabilities anchor these dynamic predictions in a stable world. Temporal coherence uses physics-aware generation to maintain view consistency, so that when a camera pans away and returns, the 3D environment remains unchanged, a persistent world state rather than a fleeting video. Feed-forward reconstruction then enables the rapid generation of 3D-native assets such as depth maps, normals, and 3DGS, which can be immediately exported to simulation platforms and rendered at 24 FPS on consumer-grade hardware.

The result is a move away from "pixels as video" and toward "pixels as world states": a foundational layer where developers can post-train generalist models on task-specific robotics data, cutting development cycles from years to days.

12.1.5 The Enterprise World Model: Applications in Business and Operations

In the corporate sphere, World Models introduce what might be called "Starcraft for CEOs." With a live operational model of the business, leaders stop relying on retrospective reporting and start running predictive simulations instead.

Robotics and Embodied AI. World Models attack the "Sim-to-Real" gap by giving agents like SIMA 2 an "unlimited curriculum." A robot can master a difficult task in a safe, high-fidelity virtual environment where the physics are indistinguishable from reality, then carry that skill into varied industrial settings without manual reprogramming.

Autopilot and Logistics. For autonomous systems, World Models enable "counterfactual reasoning." A system can ask, "What if that cyclist swerves?" and simulate 10,000 variations of that edge case in milliseconds. This move from "local capability" to "verifiably safe" navigation is the key to scaling AV fleets in complex urban terrains.

Supply Chain and Operations (Digital Twin 2.0). Following Rohit Krishnan's vision, the "Enterprise World Model" treats the business as a dynamic environment that can be rehearsed before it is run. In a real estate context, this plays out across several intertwined decisions. For capital allocation, a model can simulate the ROI of a $60k roof repair against the risk of a $500k total capex failure four months later, surfacing trade-offs that a static spreadsheet would obscure. For operational efficiency, managers can simulate the revenue lift of enforcing a 15-minute response-time SLA versus the increased staffing cost it implies, finding the Goldilocks zone for conversion. And for market strategy, companies can model price-matching responses to competitors and watch the impact ripple through occupancy and margin before the first dollar is spent.

12.1.6 Limitations, Responsibility, and the Horizon of AGI

The transition to a simulation-governed world brings ethical and technical risks at the system level. The stakes change too. We are no longer managing data errors; we are managing physical safety.

Technical Constraints and Constraints of the "Dream." Even advanced models like Genie 3 face hurdles that bound what the simulation can faithfully represent. The action space available to an agent inside these worlds is still narrow, restricting the range of physical behaviors that can be rehearsed. Temporal consistency is another frontier: maintaining high-fidelity coherence over hours or days, rather than minutes, remains unsolved, and a world that drifts over time cannot be trusted for long-horizon planning. Finally, text and geography lag behind: legible signage and accurate geographic detail in generated environments fall well short of the visual quality of the rest of the scene.

The "Simulation Gap" and Hallucinations. "World Model hallucinations" are a far more dangerous liability than LLM fact errors. If an LLM hallucinates, a sentence is wrong. If a World Model hallucinates gravity or object mass, a robot damages property or injures someone. This "Sim-to-Real" gap makes alignment a matter of physical liability, which is why rigorous "reward modules" are needed to track and penalize physically impossible outcomes.

Future Outlook: The Sensory Cortex of AGI. The AI "grandmasters" remain divided. Yann LeCun views current LLMs as a potential "dead end" for true autonomy, a conviction he backed with his career: in late 2025 he left Meta, where he had been chief AI scientist since 2013, to found a startup dedicated to world models, the clearest possible signal of where he believes the next paradigm lies. Others see LLMs as the linguistic interface for a more powerful sensory engine. The consensus, however, is that World Models will serve as the "sensory cortex" for a future AGI, giving it the ability to not just talk, but to navigate, anticipate, and exist.

Digital systems are beginning to navigate our reality, not merely process our information. The companies that build and control these internal simulators will hold a rare advantage: they can rehearse the future before it happens.

12.2 The Dawn of Agentic Commerce

12.2.1 Introduction: The Shift to Delegated Shopping

Agentic commerce hands the act of shopping to AI agents that work on our behalf, and research by McKinsey & Company suggests the shift is moving from concept to strategy. McKinsey projects that by 2030, the US B2C retail market could orchestrate up to $1 trillion in revenue from agentic commerce, with global projections reaching $3 trillion to $5 trillion. The change unbundles the traditional shopping trip. Instead of visiting vertical destinations such as Amazon or Expedia, consumers state an intent, and a personal agent acts as a concierge across a horizontal, integrated ecosystem to fulfill it.

12.2.2 The Mechanics of Agentic Commerce

According to McKinsey's analysis, agentic commerce currently takes shape across three primary interaction models.

Agent to site: Personal AI agents interact directly with merchant platforms, such as a travel agent scanning hotel websites to find and book options that fit user preferences.

Agent to agent: AI agents autonomously transact and negotiate with vendor AI agents, such as a personal shopping agent securing a bundle discount directly from a retailer's in-house AI.

Brokered agent to site: Intermediary systems facilitate multi-agent and multi-platform interactions, like a personal agent connecting with an OpenTable broker agent to secure a reservation and apply loyalty discounts.

McKinsey highlights that this ecosystem depends on infrastructure that is still being built. Three enablers stand out: the Model Context Protocol (MCP), an interoperability standard allowing agents to share context and intent across tools; the Agent-to-Agent (A2A) Protocol, which lets autonomous agents securely negotiate and coordinate across platforms; and the Agent Payments Protocol (AP2), an open standard allowing agents to make verifiable, cryptographically signed purchases. Competing with the Google-led AP2 is the Agentic Commerce Protocol (ACP), developed by OpenAI and Stripe and launched with Instant Checkout inside ChatGPT in late 2025, while Visa and Mastercard have both introduced agentic payment credentials. A protocol war is underway, and merchants are not neutral bystanders: Amazon has blocked third-party shopping agents from its storefront, and Perplexity's agent was the subject of high-profile litigation over exactly this question. Who owns the customer relationship when an agent does the shopping is the commercial fight of the coming years, and the protocols are its weapons.

12.2.3 The Automation Curve

McKinsey emphasizes that agentic commerce does not arrive in a single leap to full autonomy. It unfolds along a six-level agentic commerce automation curve, based on how much of the journey consumers are willing to delegate. McKinsey defines the levels as follows:

Level 0: Programmed convenience. The pre-agentic baseline, consisting of rules-based, "set it and forget it" subscriptions, such as recurring coffee bean deliveries.

Level 1: Assist (the cognitive sidekick). Agents gather and present information by scanning catalogs and summarizing trade-offs, but the human shopper retains all assembly and execution duties.

Level 2: Assemble (the personal shopper). Agents orchestrate purchase-ready baskets by resolving constraints (balancing price, compatibility, and delivery speed), but stop short of buying until they receive human approval.

Level 3: Authorize (the supervised executor). Consumers delegate rules and set clear guardrails (for example, "buy my preferred sneakers if they drop below $80"), and the agent executes the workflow end-to-end within those boundaries.

Level 4: Autonomize (the intent steward). Agents operate against standing, long-term goals such as maintaining airline loyalty status at the lowest cost or keeping household essentials under a monthly budget, acting proactively with only episodic human intervention.

Level 5: Networked autonomy (multiagent commerce). A multi-agent world where personal agents autonomously negotiate directly with specialized vendor agents across pricing, logistics, and payments with minimal human input.

McKinsey research notes that this curve also applies to B2B commerce, with the caveat that B2B delegation is institutional rather than personal; autonomy advances more slowly due to strict procurement policies and compliance rules, but scales far more powerfully once it takes hold.

12.2.4 Implications for Merchants and Value Pools

McKinsey's analysis shows agentic commerce collapsing the traditional sales funnel: search, comparison, and consideration merge into a single agent-mediated moment. For retailers, the competition moves away from front-end brand storytelling and toward operational trust and delivered value. Here is the hard part. If a retailer's catalogs, pricing, shipping promises, and substitution policies are not exposed cleanly through machine-readable APIs, McKinsey warns that AI agents will simply bypass them.

Because agents bypass traditional ad channels, McKinsey notes that retail media networks reliant on ad-based models face a decline in revenue. To survive, businesses will need to adopt new monetization models identified in McKinsey's research, such as charging real-time negotiation fees, facilitating multi-brand bundle revenue sharing, offering premium subscriptions for specialized vertical agents, or monetizing anonymized agent-filtered consumer analytics.

12.2.5 Navigating Trust and Risk

McKinsey stresses that in this model, trust stops being a marketing asset and becomes infrastructure. Building it comes down to two things: explainability, so consumers understand why an agent made a specific choice, and reversibility, so cancellations and human overrides are easy.

Agentic commerce also introduces new systemic risks. Payments, compliance, and fraud systems built to stop bots must learn to verify legitimate agents instead, through a "Know Your Agent" (KYA) standard. McKinsey's studies emphasize that the industry needs new accountability protocols to prevent failures where a single faulty prompt sets off a chain of unintended autonomous errors across interconnected systems.

12.3 The Agentic Enterprise: When Software Becomes a Coworker

Something changed in how executives talk about AI agents, and the survey data caught it. In the 2025 annual research report from MIT Sloan Management Review and BCG, 76 percent of executives said they view agentic AI more as a coworker than as a tool (Ransbotham et al., 2025). That is not a figure of speech. It reflects systems that take over a routine step in one workflow, support a human expert with analysis in another, and coordinate with other agents in a third, shifting decision-making authority as they go. The adoption curve is steeper than anything the field has produced before. Traditional AI took roughly eight years to reach 72 percent adoption. Generative AI hit 70 percent in three. Agentic AI reached 35 percent adoption within about a year of becoming commercially practical, with another 44 percent of organizations planning deployment (MIT SMR and BCG, 2025).

The consultancies see the same acceleration with more granularity. McKinsey's late-2025 survey found 62 percent of organizations experimenting with agents and 23 percent scaling them in at least one function, though no single function crossed 10 percent scaled deployment (McKinsey, 2025). To move past that plateau, McKinsey proposes what it calls the agentic AI mesh: a composable, vendor-agnostic architecture in which agents from different providers collaborate across systems. The firm's blunter observation is the one I would underline: unlike earlier GenAI tools that could be plugged into existing workflows, agents demand a rethinking of the business process itself.

Where the redesign happens, the numbers are striking. McKinsey documents a research firm whose multiagent data quality system projects savings above 3 million dollars a year with productivity gains near 60 percent, a major bank whose human and AI digital factory cut legacy application modernization effort by more than half, and customer service architectures that resolve up to 80 percent of common incidents autonomously (McKinsey, 2025). At the macro scale, the McKinsey Global Institute estimates that agents alone could perform tasks occupying 44 percent of current US work hours, and puts the midpoint economic value of automation for the US at around 2.9 trillion dollars annually by 2030. Its framing of the destination is worth quoting: work in the future will be a partnership between people, agents, and robots, all powered by AI (MGI, 2025).

Now the cold water. MIT Technology Review named 2025 the year of the great AI hype correction, and the evidence it assembled deserves a place next to the projections. A July 2025 MIT study found that 95 percent of businesses attempting formal AI pilots saw no measurable value within six months. A November 2025 Upwork study found that agents from the leading labs failed to complete many straightforward workplace tasks without human help (Heaven, 2025). Even Ilya Sutskever, among the strongest believers in the scaling path, conceded that current models generalize dramatically worse than people. The gap between a demo and a dependable coworker is still wide, and companies that plan as if it were closed are the ones funding the failure statistics.

Both things are true at once, which is why executives need a decision rule rather than a mood. Here is the one this book recommends: deploy today where work is frequent, verifiable, and tolerant of retries (the loop-engineering test from Chapter 3), pilot with evals where verification can be built, and treat everything else, including the boldest projections below, as an option to monitor quarterly, not a plan to fund. The $2.9 trillion and the 95-percent-failure statistics describe the same world: enormous value available precisely to the organizations that can tell which of their workflows pass that test. And the frontier keeps moving underneath the argument. OpenAI has stated publicly that it aims to field an autonomous AI research intern by late 2026 and a fully automated researcher by 2028, and frontier systems have already contributed solutions to previously unsolved mathematics problems (MIT Technology Review, 2026). Whether or not those dates hold, the direction is unambiguous: agents are moving from executing defined tasks toward pursuing open-ended goals.

That trajectory makes management, not technology, the binding constraint. An MIT SMR expert panel found 69 percent agreement that agentic AI requires genuinely new management approaches: oversight across the agent's whole life cycle, explicit human accountability for outcomes, predefined rules for when an agent's decision should prevail over a human's, and attention to the unsettling case of agents created by other agents (Renieris et al., 2025). California Management Review describes the operational shift as moving from human-in-the-loop, where a person approves every action, to human-on-the-loop, where people set boundaries and agents act freely within them (Saini, 2026). The early adopters show what that looks like in production: Lemonade settles roughly a third of insurance claims autonomously, some in as little as three seconds, and Maersk reports a 23 percent reduction in fuel consumption from autonomous vessel routing coordination.

The governance gap is the number that should worry boards. Deloitte's 2026 State of AI research found that about 74 percent of companies plan to deploy agentic AI within two years, while only 21 percent report a mature governance model for autonomous systems. We will return to what that governance should look like in Chapter 13. The point here is simpler. The Five A's framework in Chapter 7 ends with Agents for a reason. The organizations treating that final stage as an org-design problem are pulling away from the ones treating it as a procurement decision.

Discussion Questions

  1. If autonomous agents mediated 30 percent of purchases in your category, who captures the margin, and what would you start building now to be that party?
  2. Which monitor-quarterly bet in this chapter would most change your strategy if it matured two years early, and what is your tripwire for noticing?
  3. What would a world-model or simulation use case look like in your operations, and what data would it require that you do or do not have?
This chapter is part of GenAI for Business, free to read in full. Continue with the next chapter, browse the glossary, or use the free templates it references.
Generative AI for Business Model InnovationNavigating the Ethical Landscape