TwelveLabs raises $100m to turn video archives into agentic AI

A $100m Series B backs TwelveLabs' push to make video as machine-readable as text, with Amazon joining as both investor and cloud partner.

Three graphics cards with cooling fans and connected white cables sit on a table in a brightly lit laboratory, featuring glowing blue LED indicators and blurred glassware in the background.

TwelveLabs, a San Francisco and Seoul-based video intelligence company, has closed a $100 million Series B co-led by NEA and NAVER Ventures, with Amazon, Index Ventures, Radical Ventures and a clutch of others participating. The round funds a strategic pivot from selling discrete video-understanding models to delivering what the company describes as a full-stack agentic video intelligence system, one that unifies perception, knowledge retrieval and reasoning within a single, persistent architecture.

The timing is significant. Enterprises across media, government, sport, automotive and security are sitting on vast stores of footage that existing AI tools cannot effectively interrogate. The company says video accounts for the majority of global data, it cites figures ranging from 80% to 90% in different parts of its release, yet the analytical infrastructure that has transformed text over the past decade has no equivalent for moving images. TwelveLabs is positioning itself as the company that closes that gap at production scale.

From models to memory

At the core of TwelveLabs' technical architecture are two proprietary models. Marengo 3.0 functions as a video embedding engine, converting raw footage into a semantic layer that machines can search across sound, speech and motion simultaneously. Pegasus 1.5 then transforms that embedded content into structured data, scene boundaries, named entities, temporal segments, making it parseable by any intelligence layer built on top. The company compares Pegasus's function to what markup languages do for documents in a web browser: turning raw content into something a reasoning system can act on.

The new agentic layer built on top of these models is what distinguishes this funding round from a straightforward model-scaling play. Rather than querying video from scratch with each request, the system builds a persistent, indexed memory of ingested content. The company's argument is that intelligence compounds over time: the more footage the system processes, the more capable its reasoning becomes across that corpus. Its first application-layer product, Rodeo, launched earlier this month as the initial step in moving from API infrastructure to end-user tools.

Convergence read-across: compute, cloud and the agentic stack

Amazon's participation is more than a financial footnote. AWS is named as TwelveLabs' preferred cloud provider under a multiyear commitment, and future TwelveLabs models will launch first on Amazon Bedrock. Critically, TwelveLabs' video inference workloads will be optimised for AWS Trainium chips, Amazon's purpose-built AI accelerator, rather than the NVIDIA GPUs that dominate most foundation-model infrastructure. That alignment has implications beyond a single vendor relationship: it signals that the video-intelligence workload, which is computationally distinct from large language model inference, may become a meaningful driver of demand for alternative AI chip architectures.

For cross-sector investors, the strategic geography of this deal is worth noting. NAVER Ventures is the Korean internet giant's venture arm; Korea Investment Partners also participated. TwelveLabs operates engineering teams in both San Francisco and Seoul, and the company is opening offices in New York and London to serve enterprise demand. That Korea-US axis, combining frontier AI research with deep pools of media and government procurement, mirrors the dual-market model now common in semiconductor and defence-adjacent AI companies.

The vertical reach is deliberately broad. TwelveLabs cites media and entertainment as its deepest traction, but government and public-sector contracts are growing: the company says it is working with administrations globally on mission-critical workflows. Security and automotive are named as further growth vectors. That range matters for capital allocators assessing whether video AI is a point solution or a genuine horizontal infrastructure play.

The open question is competitive durability. Google DeepMind, Meta and OpenAI are all building multimodal capabilities that increasingly touch video. TwelveLabs' counter-argument, that models commoditise but the intelligence layer that orchestrates them does not, is a recognisable thesis in the foundation-model era. Whether the company's end-to-end ownership of perception, knowledge and reasoning creates a moat durable enough to hold as hyperscaler multimodal capabilities mature will define the next chapter of its $100 million mandate.