For years, the AI industry has been chasing bigger context windows. More tokens. More memory. More information fed into models all at once. But behind the race to build smarter systems sat a stubborn mathematical problem that every major AI lab on Earth was forced to live with: the larger the context window became, the more painfully expensive the system became to run.
Now, Subquadratic says it has found a way around that constraint.
The Miami-based startup emerged from stealth this week with the claim that it has built the first large language model with fully subquadratic scaling, meaning compute grows linearly with context length rather than exponentially. If validated independently, the breakthrough could dramatically lower the cost of processing massive amounts of information through AI systems.
The company’s first model, SubQ 1M-Preview, reportedly reduces attention compute by nearly 1,000 times at 12 million tokens compared with standard transformer architectures. Alongside the launch, the company unveiled a coding product called SubQ Code, a search tool called SubQ Search, and an API currently available through private beta.
The launch quickly spread across the AI world. According to Co-founder and CEO Justin Dangel, the announcement generated more than 12 million views on X and over 30,000 waitlist signups within the first 24 hours.
The excitement surrounding the company stems from a real technical limitation that has shaped the economics of modern AI. Nearly every frontier model today relies on transformer architecture, where every token compares itself against every other token inside a sequence.
“The entire LLM world is built on transformers,” Dangel [pictured above] told Refresh Miami. “The architecture has a limitation. When you’re processing paragraphs and paragraphs of information, it becomes too expensive to handle large amounts of data at once, even at the frontier level.”
That limitation has shaped how developers build AI applications today. Instead of feeding entire datasets into models, teams often rely on retrieval systems, vector databases, prompt engineering, chunking systems, and orchestration layers to narrow information before it ever reaches the model.
Subquadratic believes much of that complexity exists because the underlying architecture cannot efficiently process long contexts.
The company’s approach, called Sparse Subquadratic Attention, selectively focuses only on the token comparisons that matter rather than computing attention across every possible relationship. In theory, that dramatically lowers computational overhead while still preserving retrieval quality across extremely large contexts.
The company says the architecture allows it to expand context windows while operating at significantly lower cost. “It allows us to operate much less expensively,” Dangel said.
Still, the launch has triggered sharp debate across the AI research community.
While some researchers described the work as potentially significant, others questioned whether the company’s claims fully hold up under scrutiny. AI commentator Dan McAteer summarized the polarized reaction in a widely shared post: “SubQ is either the biggest breakthrough since the Transformer… or it’s AI Theranos.”
Dangel said he understands the reaction. “Extraordinary claims will often be greeted rightly with skepticism,” he asserted. “The fact that our company has a potentially industry-disrupting innovation, I’m not surprised by the reaction.”
He added that the company plans to release additional technical papers and products in the coming months. “We look forward to releasing products and papers,” he said. “I hope the community will be satisfied.”
Subquadratic has raised $29 million in seed funding from investors including Tinder co-founder Justin Mateen, #MiamiTech stalwart Javier Villamizar, and investors previously involved with companies including Anthropic, OpenAI, Stripe, and Brex.
According to Dangel, who invested in Subquadratic before joining as CEO, the company has been building the technology for roughly five years.
Previously, co-founder and CTO Alexander Whedon worked as a software engineer at Meta and later served as Head of Generative AI at TribeAI. Dangel said the broader research team includes 11 PhD researchers from organizations including Meta, Google, Oxford, Cambridge, ByteDance, Adobe, and Microsoft.
The company is also planting its flag firmly in Miami. “Miami is the best place to live in the world,” Dangel said. “Miami is a real innovation center. This is a great place to build if you have the right team.”
READ MORE IN REFRESH MIAMI:
- Hovhannes Avoyan was doing AI before it was cool, and now Miami is home base for his next act
- Everyone bought AI tools. Certifyde wants to make people actually use them – and just raised $2M
- eMerge day 2 cut through the noise on AI, startups, and what it really takes to scale in South Florida
- Cast AI is tackling the $30 million problem sitting idle in your cloud
- Koywe’s answer to slow international payments started with a lifelong friendship - September 9, 2026
- Split Pay raises $125M to give your paycheck better timing - September 9, 2026
- PRIMA raises $5M to turn hot tables into a referral business - September 4, 2026
