Erica · PM Research MemoAll memos

PM RESEARCH MEMO

Title: Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology
Author / source: Naveen Rao (co-founder/CEO, Unconventional AI); All-In Podcast stage; Chamath Palihapitiya joins Q&A (~19:40+)
Source title: Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology
Source URL: https://www.youtube.com/watch?v=yAsrMA_ADPc
Video ID: yAsrMA_ADPc
Channel: All-In Podcast (@allin)
Published / upload: 20260921 (Mon Sep 21, 2026) — desk flag ~5:04 PM ET per brief
Duration: 22:34 (1354s)
Memo date: Tuesday, 22 September 2026 (America/Toronto)
Views / likes at pull: ~63,465 / ~646 (~39 comments)
Transcript: /workspace/youtube-transcripts/yAsrMA_ADPc.md · Brief: /workspace/youtube-transcripts/yAsrMA_ADPc_brief.md · Meta: /workspace/youtube-transcripts/yAsrMA_ADPc.json
Caption source: YouTube ASR only (en-orig json3 via yt-dlp). No manual English captions. Apply ASR locks below; do not quote raw ASR proper nouns without cleanup.
Source type: Conference / All-In stage talk + Chamath product Q&A. Motivated founder (Unconventional AI CEO pitching a new compute substrate). Channel / format: opinion tech podcast / conference stage — not investment advice.
Product: Institutional PM research map of claims on AI energy wall, von Neumann inefficiency vs biology, dynamical / “4D” computing, Unconventional AI prototype/roadmap, Jevons-scale demand, and product path (rack/DC system). Not advice. No buy/sell from this desk. Language such as “anti-doomer hardware bull / energy-wall / 1000× / Jevons” is source expression / research hypothesis, not an Erica/desk recommendation.
Source discipline: Primary = this transcript + brief + meta JSON only. Companions (Jensen All-In S7CrlFLAmEA, Asianometry x8QP9oXgahA) may be cited only as External / other desk for cross-reference — do not silently import their numbers into underwriting. Flag any other fact as External check needed.

ASR name / token locks (from brief):

ASR / garble Intended / note
Nirvana systems Nervana Systems (sold to Intel)
platform rising GPUs MosaicML (GPU infra / model training platform → Databricks)
Ali and team at data bricks Ali Ghodsi / Databricks
unconventional AI Unconventional AI (confirm branding)
map will matmul / matrix multiply
Jeban's / Jevons Jevons paradox
VM that sits somewhere Product framed as managed rack/DC appliance (not classic VM) — Chamath prompt; Naveen reframes as DC rack/system
Uno Image-gen model on oscillators (as named) — confirm public write-up External

Sponsor noise (ignore as claims): IREN, Oracle, EY, Meta, Keel Infrastructure, Airwallex, PayPal, Google for Startups (and cold-open montage).

Guest pedigree (as stated on tape / brief): Founded first AI-chip co Nervana (2014) → sold to Intel, ran Intel AI group; then GPU infra / model-platform co (MosaicML) → joined Databricks (2023); Mosaic path now “~¼ of Databricks total revenue” (as-spoken); now Unconventional AI (rethink computer for power efficiency).


EXECUTIVE SUMMARY


SOURCE-ACCURATE SUMMARY

Chapter-ordered from official description / meta JSON. Quotes ≤20 words. Attribution by speaker. Timestamps approximate (ASR).

Welcome Naveen Rao (0:00–4:45)

  1. (0:00–0:38) Intro montage / stage welcome. Hosts: Naveen Rao, co-founder/CEO Unconventional AI; “definitionally outlier founder”; built/sold two deep-tech cos; “I’m the opposite of an AI doomer”; “We need innovation on the hardware substrate.” Standing welcome. Ignore sponsor mid-rolls / cold-open music throughout.

  2. (0:38–1:12) Anti-doomer frame. Naveen: “I’m the opposite of a doomer”; AI “one of the most transformational technologies”; “next level of evolution”; All-In as “anti-doomer conference.”

  3. (1:12–1:45) Bio — early computers / EE / PhD neuroscience. Childhood computer ~1978 / early ’80s; programmed as puzzle; electrical engineer from sci-fi / intelligent-machine ambition; later PhD in neuroscience — “How do we make computers intelligent?”; world moved that direction.

  4. (1:45–2:51) Nervana → Intel; then Mosaic path. Founded “first AI chip company” ASR “Nirvana systems” → lock Nervana (2014); hard to raise when AI not vernacular; sold “way too early” to Intel; started/ran Intel AI group. Post-2020: infrastructure for bigger models / LLMs; started ASR “platform rising GPUs” → lock MosaicML; after ChatGPT 2022 “best game in town”; joined Databricks 2023 with “Ali and team” → Ali Ghodsi.

  5. (2:51–3:24) Mosaic economics claim; Unconventional pitch. Mosaic path “actually that’s a quarter of the total revenue of data bricks today” — as-spoken. Unconventional AI: “rethinking the foundations of how a computer works”; singular purpose power efficiency.

  6. (3:24–3:57) 1000× timeline revised. Initially within 5 years to 1000× power efficiency; revised to three and a half years — “things have gone faster”; “solved very deep scientific problems quicker because of AI.”

  7. (3:57–4:45) Org stack top-to-bottom. Theorists (math PhDs, theoretical neuroscience) → concepts that move less information → models trained/evaluated on real data → physical circuit architects/designers → board/system/product. Segue: “So, is energy really a problem?”

Is energy really the problem? (4:45–10:24)

  1. (4:45–5:37) Google tokens → GW math. One company (Google, public): 3.2 quadrillion tokens/month. Assume 10 joules/token (low end) → 12 gigawatts. US puts ~40 GW into data centers; US ~half world DC capacity → world under 100 GW. “12 gigawatts is going into one company just for AI services.”

  2. (5:37–6:13) Run-out / gap graphic. Models bigger + demand growing → energy up; “run out of energy pretty fast, in like 3 years or so.” Exponential AI market vs linearized energy: call it trillion-dollar market in 2030, maybe bigger; gap is the problem; solve with technology.

  3. (6:13–6:46) Scarce input shift; 50% energy. DC thinking: floor space → networking → GPUs → energy first (get power contract, then fill with GPUs). “About 50% of the cost of serving a token… is energy”; rest CapEx/hardware/floor.

  4. (6:46–7:18) Watt monetization business case; brain 20 W. Get power contract → monetize every watt; “monetize that 1,000x better than existing hardware.” Biology proof: human brain ~20 watts.

  5. (7:18–7:51) Animal brains / phone comparison. Monkey-scale ~1 W (phone ~1 W); rats/bats milliwatts; squirrel brain eight milliwatts; “over a hundred squirrel brains on your phone”; precise motor behavior.

  6. (7:51–8:55) Understand by creating; bit-rate contrast. Quote frame: don’t truly understand until we can create it. Most energy in computing = moving information. Cortex ~16 billion bits/s (~13–14B neurons); GPU/high-end ~30 trillion bits/s memory in/out; inside chip 10–100× more — drives energy demand.

  7. (8:55–10:24) History of computers; ENIAC; Moore end. Mechanical → analog → digital 1930s–40s; operation similar today — memory outside, compute, shuttle bits; built for speed not efficiency. ENIAC (1945) artillery trajectories faster than human computers; today sell “twice the speed.” Transistor counts up, frequency/single-thread stalled; efficiency from smaller transistors “largely ended” (Moore). Need rethink: “cut out the middle man.”

Cutting out the middleman (10:24–19:40)

  1. (10:24–11:07) Abstractions are lossy. Digital 0/1 is abstraction of transistor continuum; stacked abstractions → neural nets on top; each lossy. Simplify: abstraction of semiconductor physics connected to neural network. Brain: neurons, “no linear algebra… no floating point math”; physics gives rise to intelligence — mimic with semiconductor.

  2. (11:07–12:14) Dynamical systems in nature. Computation in nature: bird flocks, ant colonies; dynamical systems theory — emergent properties from simple component rules; brain works this way; building circuits from these ideas. Metronome example: multiple metronomes on rolling plank synchronize via physics.

  3. (12:14–13:54) Metronomes → Uno image model. Can such a system do computation for generative AI? Released model Uno — image generation on oscillators; simulated; open-sourced as claimed; state-space trajectories conditioned on class (airplane/car/bird); actual generated images shown.

  4. (13:54–15:33) Sparsity as holy grail. Dense all-to-all = n² connections (10→100; 1000→1M). Sparsity: throw away connections, rescue (even improve) behavior; more trainable; works in simulation and physical systems; “more efficient… more scalable… more performance.”

  5. (15:33–16:06) First physical dynamical computer. “First time… talking about this publicly”; “first physical dynamical computer ever built”; company earnest January (no team initially); taped out June 1; chip back in lab with results; first images from such a computer. Applause.

  6. (16:06–16:39) 500 nJ/image claim. Beyond images: sequence modeling / language models possible as claimed; ~500 nanojoules per image vs GPU millijoules order; “many orders of magnitude more efficient”; “doesn’t move information around”; “proof positive.”

  7. (16:39–17:46) Von Neumann vs dynamical; 4D computing. CPU→GPU→compute-in-memory still von Neumann (memory↔compute shuttle). Dynamical computer: compute and memory co-located; no memory interface; each element is memory. 4D computing: time/dynamics + physical 3D die stacking (vertical + planar).

  8. (17:46–18:52) Thermodynamic limit; beat biology; local DCs / robots. Intelligence per watt; thermodynamic limit never exceedable; mammalian brains within 1–2 orders; today ~10 billion× away. In 3.5 years hit limits of 2D lithography; company goal beat biology this decade → compute everywhere / robotic forms. Shift: big GW campuses → many small local DCs; environmentally friendlier, local, adaptive; enable billions of robots.

  9. (18:52–19:40) Jevons paradox close. AI ~trillion-dollar market; disrupt by 1000× → Jevons paradox (ASR “Jeban's”): drop cost → consume more than the drop; 1000× cheaper → consume more than 1000×; “largest market that humanity’s ever seen.” Applause; Chamath: “extremely unexpected.”

Chamath joins: path to product (19:40–22:34)

  1. (19:40–20:29) Ecosystem / product / timing. Chamath: need fabs, packagers, ecosystem beside you (nod to Jensen earlier — External companion context only); path from early version to something people use? Naveen: within 2 years to full product. Product = new data center product first — whole rack/system; tokens in/out through network cable; “inner guts… completely different.” Chamath “VM” prompt → Naveen reframes as managed DC rack/system (ASR lock).

  2. (20:29–21:35) Porting models; matmul. Chamath: existing model families / KV cache / reductive abstractions? Naveen: sliding scale of better vs pain; make move compelling; port at model layer not operations layer; existing models will work but “fair bit of compute” for transition. Matmul (ASR “map will”): can characterize as matmul but doesn’t implement as matmul; time-varying behavior; each timestep analyzable as state × transition matrix.

  3. (21:35–22:34) Team span; Python libs. Team = dynamical-systems theorists (field ~100 years) + chip builders — “they don’t talk to each other”; facilitating that span “one of the most challenging things.” CUDA analogy denied: Python libraries expressing time-varying stochastic elements. Chamath close: ambitious; thanks.


SYSTEMS MAP / VALUE CHAIN

Source-locked chain (power → DC → substrate → models → apps/robots):

  1. Power / energy contracts (scarce input). As-spoken: today’s DC scarce input is the energy contract first, then fill with GPUs/infra. ~50% of token serving cost = energy. US ~40 GW DC / world <~100 GW; Google AI alone framed at ~12 GW under 10 J/token. Binding constraint narrative: ~3-year energy wall if growth continues.

  2. Data-center shell / networking / floor (demoted but still necessary). Historical scarce inputs: floor → networking → GPUs. Still needed to host racks; Unconventional first product is a whole rack/system sitting in that shell with tokens over network.

  3. Incumbent compute substrate (GPU / von Neumann stack). CPUs → GPUs → compute-in-memory — still memory↔compute shuttling; optimized for speed; energy dominated by moving bits (~30T bits/s memory I/O vs cortex ~16B). Pedigree contrast: Nervana (early AI chip) → MosaicML (GPU scale infra) → now escape that substrate.

  4. Unconventional dynamical / 4D substrate (claimed alternative). Oscillator / dynamical-systems circuits; sparsity; compute+memory co-located; time dimension + 3D stacking. Prototype silicon taped out Jun 1; ~500 nJ/image claim. Org: theorists → models → circuits → boards → product.

  5. Model layer / software port. Not CUDA; Python libraries for time-varying stochastic elements. Port at model layer (compute-heavy transition); no classic matmul implementation (state × transition characterization). Existing model families claimed workable after port — ecosystem attach risk sits here.

  6. Serving economics / watt monetization. Business case: monetize each watt 1000× better. If true, power-contract holders + Unconventional attach capture surplus; if false or late, GPU stack continues to clear scarce watts at lower intelligence/watt.

  7. Deployment topology (long horizon). Source vision: from giant GW campuses → many small local DCs; enable billions of robots; “beat biology” decade goal; compute everywhere.

  8. Demand multiplier (Jevons). 1000× cheaper intelligence → consume >1000× → “largest market humanity’s ever seen.” Efficiency success ≠ demand destruction in this narrative.

[Energy / power contracts] → [DC shell / network]
        ↓
[Incumbent GPU / von Neumann]  ↔  [Dynamical / 4D Unconventional rack]
        ↓
[Model-layer port / Python libs]
        ↓
[Token serving / agents] → [Local DCs / robots / “compute everywhere”]
        ↓
[Jevons: more total intelligence demand]

Incentive note: Founder of Unconventional AI has every reason to narrate an imminent energy wall and a 1000× substrate breakthrough. Pedigree (Nervana, Intel AI, MosaicML, Databricks revenue claim) adds credibility and sales skill. Treat as motivated but coherent; separate claim classes (public Google token figure vs private nJ measurements vs timeline promises).


SECOND AND THIRD-ORDER EFFECTS

Causal chains. S = source claim; I = desk inference (hypothesis). Companion citations = External / other desk only — no number import.

Chain A — Energy wall binds (~3y) → power becomes the allocation clock

Chain B — Credible 100×–1000× substrate → who monetizes the watt?

Chain C — Model-layer port + no matmul → software moat / switching-cost theater

Chain D — Jevons success → more total demand, different topology

Chain E — Motivated anti-doomer stage → narrative coupling with Jensen All-In week


SCENARIO FRAMEWORK

Probabilities are desk research hypotheses for organizing diligence — not forecasts or trade tickets. Motivated founder speaker.

Bull — “Wall binds; dynamical rack attaches; Jevons fires” (~25%)

Assumptions: Energy wall narrative confirmed by hyperscaler power constraints within ~3y; Unconventional (or peer dynamical/analog approaches) ships rack product with verified orders-of-magnitude energy/token gains; model-layer ports for at least one production family complete; power-contract holders adopt alternative racks for inference slices; Jevons lifts total token/robot demand.
Winners (hypothesis): Unconventional / similar substrate startups (private); power owners who re-monetize watts at higher intelligence/MW; local/micro-DC and robot-adjacent infra; selective packaging/3D-stack suppliers.
Losers (hypothesis): Pure “GPU CapEx forever linear” theses that ignore watt scarcity; inefficient serving stacks; late NeoClouds without power or efficiency path.
Leading indicators: third-party nJ or J/token audits; named hyperscaler/NeoCloud pilots; 2y product ship evidence; rising intelligence-per-MW disclosures.

Base — “Energy constraint real; substrate replacement slow” (~50%)

Assumptions: Power remains scarce and ~token energy share stays material; GPU/packaging efficiency and software continue incremental gains; Unconventional prototype is scientifically interesting but product attach slips past 2y; model-port friction high; 1000×/3.5y misses but 10–100× research progress continues; Jevons mostly shows up as more GPU tokens, not rack displacement.
Winners: Diversified AI infra (power, interconnect, incremental GPU efficiency); hyperscalers with locked PPAs; tooling that reduces bits moved (sparsity, better serving).
Losers: Narrative-only “beat biology” equity stories; CapEx without energization; pure pause theses that ignore demand.
Leading indicators: utilization vs energized MW; Unconventional fundraising vs missed milestones; incumbent efficiency roadmaps; token growth vs DC GW growth (External).

Bear — “Demo ≠ product; wall deferred; CUDA wins on inertia” (~25%)

Assumptions: 500 nJ/image does not generalize to frontier LLM serving; fab/yield/3D-stack or noise/precision issues block scale; theorists–chip-builder org fails; labs refuse model-layer ports; energy supply linearizes up (gas, nuclear, interconnect) enough to defer “3y wall”; founder narrative premium collapses on missed 2y product.
Winners (relative): Incumbent GPU ecosystem and CUDA software moat; utilities/developers who still energize conventional DCs; skeptics of analog/dynamical compute history.
Losers: Unconventional and peer “new machine” narratives; investors who underwrote 1000× on stage demos; any book that shorted watts assuming efficiency kills demand (Jevons fail and wall deferred is different bear).
Leading indicators: failed customer PoCs; adverse technical replications; silence after Jun tape-out hype; large new conventional GW energized on schedule (External).


COMPANY / ASSET WATCHLIST

Hypotheses for research queues — not trade tickets; no buy/sell.

Asset / cluster Thesis hook from source Metrics / catalysts to watch Risks
Unconventional AI (private) Dynamical/4D computer; 1000×/~3.5y; rack product ~2y; 500 nJ/image prototype Tape-out follow-ons, third-party energy metrics, rack LOIs, capital raises, fab partners Founder narrative; demo≠product; org-span; ASR-only claims
NVDA / GPU ecosystem Incumbent von Neumann/GPU substrate; Chamath nods Jensen/ecosystem need; bits-moved energy critique Efficiency roadmap, serving J/token, attach to power-constrained DCs Displacement if alternative racks work; also beneficiary if Jevons lifts total tokens before displacement
AMD / custom silicon (hyperscaler ASICs) Alternative paths to watt monetization without Unconventional physics Inference ASIC deployments, J/token vs GPU Same wall; different architecture bets
TSM / advanced packaging / 3D stack names “4D” uses physical 3D die stacking + time CoWoS/SoIC-class capacity, stacking yields Cycle, customer concentration; Unconventional may or may not be meaningful volume
Hyperscalers (GOOGL explicit; MSFT/AMZN/META orbit) Google 3.2Q tokens public cite; power contracts first; ~50% token cost energy Token growth, DC GW, PPA/interconnect, energy share of COGS CapEx air-pocket; custom silicon; regulation
NeoClouds / DC REITs / shell developers First Unconventional product = DC rack/system; topology → many small local DCs long-term Energized MW, utilization, willingness to host non-GPU racks Financing; stranded shells; GPU-only designs
Utilities / IPPs / nuclear / gas peakers Energy wall ~3y; watt monetization; Jevons may raise aggregate demand Interconnection queues, AI offtake PPAs, local-DC load growth Policy, permitting, overbuild
Databricks (private) / Mosaic path Pedigree + “~¼ of Databricks total revenue” as-spoken Mosaic contribution verify (External); talent spinouts Founder claim inflation
Intel (pedigree only) Nervana sale; Naveen ran Intel AI group Historical context only — not a thesis from this tape N/A from primary
Robot / edge compute OEMs (long-dated) Billions of robots; beat biology; compute everywhere Only after rack credibility — watch as optionality Decade vision; timing

External / other desk cross-read only: Jensen All-In (S7CrlFLAmEA) for anti-doomer buildout / NeoCloud/power frame; Asianometry (x8QP9oXgahA) for scarcity/glut frame — structure only, no number import.


DILIGENCE QUESTIONS & RESEARCH AGENDA

  1. Verify as-spoken macros before any model input: Google 3.2 quadrillion tokens/month; ~10 J/token → ~12 GW; US DC ~40 GW; world <~100 GW; ~50% token serving cost = energy; ~$1T 2030 AI market frame; brain 20 W / squirrel 8 mW / cortex 16B bits/s / GPU 30T bits/s. All → External check needed (units, scope: training vs inference, PUE, etc.).
  2. 500 nJ/image: task definition, image size/model, measurement method (chip-only vs system), comparison GPU baseline SKU — demand third-party or paper replication.
  3. Uno model: locate open-source write-up / weights; what task quality vs diffusion baselines; simulation vs silicon gap.
  4. Jun 1 tape-out: process node, foundry, package, yields, how many dies, what “results” include beyond demo images.
  5. 1000× in ~3.5y / full product ~2y: 1000× against which baseline (which GPU, which workload — image vs LLM token)? Product = inference-only rack or training too?
  6. Model-layer port: cost/time to port one production LLM or multimodal model; who pays the “fair bit of compute”; numerical parity vs approximate dynamical substitute.
  7. Matmul characterization: when is state×transition equivalent enough for frontier training dynamics? Precision, noise, stochasticity limits.
  8. Mosaic ~¼ Databricks revenue: confirm with Databricks disclosures or credible secondary — founder claim until verified.
  9. Team / capitalization: size of theorist vs chip teams; investors; exclusive fab capacity; CUDA-talent competition.
  10. Power-contract attach economics: who buys the rack (hyperscaler, NeoCloud, enterprise)? Pricing in $/token or $/watt-month?
  11. Thermodynamic / “10 billion× away” / beat biology: define metric (Landauer? intelligence benchmark?); avoid importing as tradable KPI without definition.
  12. Companion triangulation (External only): re-read Jensen All-In memo and Asianometry scarcity memo for frame conflict (buildout vs glut vs new-substrate), not number import.
  13. ASR cleanup pass: Nervana, MosaicML, Ali Ghodsi, matmul, Jevons, Unconventional branding — confirm before client-facing quotes.
  14. Legal/IP: dynamical-systems / oscillator compute prior art (decades of analog/neuromorphic attempts) — differentiation diligence.

RISK ANALYSIS

Thesis risk

Timing risk

Execution risk

External / exogenous risk

Process / desk risk


APPENDIX — Unverifiable / ASR / External check list

As-spoken numbers — External check needed before underwriting

Claim Speaker Status
Google 3.2 quadrillion tokens/month Naveen As-spoken (cites public)
~10 J/token (low end) → ~12 GW Naveen As-spoken math
US DC ~40 GW; world <~100 GW Naveen As-spoken
Energy wall ~3 years Naveen Estimate
~$1T AI market 2030 (maybe bigger) Naveen Frame
~50% token serving cost = energy Naveen As-spoken
Monetize watts 1000× better Naveen Business case
Brain ~20 W; monkey ~1 W; squirrel ~8 mW Naveen As-spoken
Cortex ~16B bits/s; ~13–14B neurons Naveen As-spoken
GPU ~30T bits/s memory I/O; 10–100× inside Naveen As-spoken
1000× in 5y → revised ~3.5y Naveen Roadmap
Full product ~2 years (DC rack/system) Naveen Roadmap
~500 nJ/image vs GPU mJ Naveen Prototype claim
~10 billion× from thermodynamic limit; brains within 1–2 orders Naveen Framing
Mosaic path ~¼ of Databricks total revenue Naveen Founder claim — verify
Company earnest Jan; tape-out Jun 1 Naveen As-spoken chronology

ASR / identity — do not cite raw

Sponsor noise — not claims

IREN, Oracle, EY, Meta, Keel Infrastructure, Airwallex, PayPal, Google for Startups

Companion memos — External / other desk only (structure cross-ref; no number import)

Primary paths


CLOSING — Packet routing

Artifact only. Markdown saved to the path above. Do not publish to here.now. Do not email. Do not auto-handoff to Compound or Jordi. Deliver in chat for AP / Erica desk packet routing. Thesis language (anti-doomer hardware bull / energy wall ~3y / 1000× dynamical-4D / Jevons demand) is source-locked research hypothesis, not a desk recommendation. Motivated founder. ASR locks enforced. No buy/sell.

— End of PM RESEARCH MEMO —

Desk copy · not a trade recommendation · Erica · 22 Sep 2026