Refiant launches 10M-token AI model targeting fintech and enterprise

A three-person UK-founded startup ships the largest publicly available context window, targeting fraud detection and long-horizon enterprise reasoning.

A cityscape featuring numerous tall, modern glass skyscrapers against a clear, bright blue sky.

Refiant AI, a California-based startup founded in 2025, has launched Protea, a suite of large language models anchored by a 10 million-token context window. The company says the flagship model can hold roughly 7.5 million words, or around 15,000 pages, in active memory simultaneously, processing them in a single pass rather than in the retrieved fragments that characterise most current enterprise AI deployments. Protea is available immediately with no waitlist, across tiers of one million, five million and ten million tokens.

The launch follows a $5 million seed round led by VoLo Earth Ventures and the announcement of research partnerships with Imperial College London and UCL's Sargent Centre for Process Systems Engineering. Refiant's three co-founders, CEO Dr Viroshan Naicker, CPO and COO Mathew Haswell, and Siddharth Gutta, span quantum mathematics, traditional finance and commercial operations. Their underlying methodology draws on evolutionary search and swarm-style optimisation, techniques modelled on how biological systems solve complex problems efficiently. An earlier application of the same approach reduced OpenAI's GPT-OSS-120B to run on a consumer-grade MacBook Pro with 18GB of RAM, a result the company credits with securing its seed funding.

Context and the "lost in the middle" problem

The dominant constraint on enterprise AI deployment today is not raw model capability but working memory. Most leading models lose coherence well before their stated token limits, a failure mode researchers call the "lost in the middle" problem: accuracy degrades sharply for information buried in the centre of long inputs while performing better at the start and end. Refiant claims Protea addresses this directly, maintaining fidelity across the full context. The company also says it has a working internal prototype at 100 million tokens and is exploring how to productionise it.

For comparison, Anthropic's Claude and OpenAI's ChatGPT currently top out at roughly one million tokens apiece, though both companies have signalled longer-context roadmaps. Refiant's claim to be first to ship a 10 million-token model in production at open access, rather than behind an enterprise approval process, is the strategically meaningful distinction here, though the "first of its kind" framing is the company's own and warrants independent verification.

The cross-sector read: fintech, clinical data, and agentic workflows

Refiant's marketing leads with fintech applications: fraud detection and credit decisioning systems that can reason across multiple years of transaction history in a single inference pass. This is a genuinely significant operational shift. Current retrieval-augmented generation architectures in financial services require chunking and retrieving transaction histories, introducing the risk of missed patterns at the seam between retrieved fragments. A sufficiently reliable long-context model removes that architectural dependency, reducing both engineering complexity and a category of model hallucination risk.

The implications extend well beyond financial services, which is where the macro capital story becomes interesting. The same long-context capability is directly applicable to clinical trial archives in biotech and pharma, multi-year regulatory dossiers in energy and infrastructure, and the sprawling codebases of enterprise software. Agentic AI systems, increasingly the dominant paradigm as enterprises look to automate multi-step operational workflows, are particularly constrained by context limits: an agent that loses track of its earlier reasoning mid-task is operationally unreliable. Protea's architecture, if the accuracy claims hold up under independent stress-testing, addresses what has been one of the most cited blockers to production-grade agentic deployment.

Investors are watching the long-context race closely. The competitive dynamic is not only between foundation-model labs but also between infrastructure-layer optimisation companies like Refiant and the hyperscalers building proprietary context-extension techniques into their own stacks. A seed-stage, three-person company claiming to leapfrog Google, Anthropic and OpenAI on a core capability metric will face intense scrutiny, but the open-access, no-waitlist strategy is a deliberate wedge: it invites the developer community to stress-test the claims before the next funding round. Refiant has signalled a three-stage product roadmap with further announcements in the following three months, suggesting the seed capital is earmarked for a Series A raise in the near term.