Cerebranalysis | Part 2: The Perils of Real World Inference
Turning experience testing the only custom designed model for Cerebras (GPT 5.3 Codex Spark) into an investment conclusion
Welcome (back) to Cerebranalysis (part 2)!
Yesterday, we covered Cerebras’ technology (dinner plate computing), which is a prerequisite to this article as we build upon the main conclusion from that piece.
Today, we will be going over the financial and business side of things and actually analyze Cerebras as an investment. This will include extensive testing of an actual model running on their hardware for real-world agentic inference tasks. Exciting right?
By accessing this content, you acknowledge and agree to our terms and conditions. This research is not financial advice.
Financials
Cerebras IPO was 20 times oversubscribed at the initial $130 share price. People really want to buy this thing. Yesterday, they raised the IPO price to $160. Polymarket still has almost a 100% chance that they will close above a $50 billion market cap (that is for their basic equity value, which is only around 70% of their fully diluted, so it is way higher than what’s shown in the table).
Their 2025 revenue is miniscule and irrelevant at $510m. Currently, they have $24.6b RPO and $20b of contract value from OpenAI, which can be expanded over time, but this implies a $6.7 billion steady-state ARR.
At the assumed IPO price of $160 per share multiplied by the fully diluted share account, the implied revenue multiple to OpenAI’s ARR is 7.4x.
The First Noble Truth of Cerebras
The following is the first Noble Truth of Cerebras.
The financials are useless. Completely utterly useless. Forward estimates are useless. Growth rates are useless. Multiples are useless.
Now, why is that? Let’s think about this from first principles.
How To Think About Cerebras
Cerebras is a canonical example of an unproven technology early in its S-curve with potential for high double-digit to triple-digit growth. This high-level conceptual profile allows us to anchor to two companies that have plenty of coverage on this channel: Lumentum and Bloom Energy.
Lumentum
The 30,000ft intellectual framing of Lumentum is simple: one market (optics) is guaranteed (by the laws of physics) to replace another (copper) to achieve full dominance of a certain use case (AI networking).
Because the transition is inevitable, the leading players in the new market have a massive guaranteed growth runway. This is why optics are so hot.
At the same time, the growth is measured and measurable. It is easy to model. We can confidently predict that Lumentum has 50% year-over-year growth because we can simply use the growth rate of optics as an anchor.
Bloom Energy
Bloom is slightly different. One company provides an objectively superior solution for one single market, which has the potential to fully displace the incumbent.
This is different from Lumentum because the market doesn’t change. Bloom plays in behind-the-meter energy, just like GE Vernova. They are direct competitors. There is no copper equivalent. There is no transition from one market to another.
Because of that, Bloom’s growth can be a little bit more unpredictable and bursty. There is no big Kager number to go off of. It is just a matter of how quickly customers adopt the new solution and how much, or if, the new solution is superior to the old one. However, growth here is still predictable. If the new solution is better, it will still slowly displace the incumbent.
Cerebras
Cerebras is completely different.
From our first article, we ended with the ultimate conclusion that Cerebras is a provider of a fundamentally different product than Nvidia: premium tokens. Premium tokens are way faster but cost way more.
Is Cerebras like Lumentum? No, there is no world in which premium tokens will inevitably displace normal tokens. They are simply two different products each with a different trade-off.
Is Cerebras like Bloom Energy? No. WSE is not a superior way to generate the same end product (electricity for Bloom, regular tokens for NVIDIA).
In both comparisons, Cerebras is actually much worse. It’s not guaranteed to take share from anyone, really. However, there is one critical distinction, and it’s very critical. Very, very critical.
The inference market is the single biggest market in the world.
AI networking and behind-the-meter power are big markets, but they’re not the biggest. Inference is, and the current incumbent is the biggest company in the world. If these premium tokens carve out a large enough slice of this overall pie, Cerebras is a very attractive investment. The only question is, will it carve up a larger left slice? Again, unlike Lumentum, it is not guaranteed to take any share at all.
Now you probably understand why I say that the financials are useless. Single-digit billions of ARR do not matter at all. If you are unprofitable and your market is the biggest market in the world, what matters much more, and really what matters at all, is the slice of the fast inference pie.
No amount of revenue will save Cerebras from offering a fundamentally unwanted solution. No valuation is too high if they serve a sizable fraction of all inference.
Therefore, the point of this article is to simply assess how likely it is the fast inference market ends up being incredibly large.
Abstract
To assess how likely it is that the fast inference market ends up being large, we must observe and test how performant and useful fast inference is today.
There is significant alpha in the testing. Just by observing the everyday experience of using the model, we can literally feel the components that bottleneck inference (hint one of the market’s current favorite themes) and reason about whether Cerebras is positioned to solve it.
For my testing, I completely disregard classic benchmarks that are unrepresentative of actual work and instead focus on two real-world tasks that I perform every day: equity research and coding. For equity research, we ask the model a classic retrieval question, and for coding, we have it develop a feature in a (real!!) live application and compare them side by side.
Through these two tests we observe what actually bottlenecks inference speed on a day-to-day basis and link this to our silicon world. We evaluate 5.3 codex Spark on both its real vs claimed speedup and the performance relative to the frontier.
We synthesized the results from our tests into a conclusion about the speed and the performance quality of 5.3 Codex Spark and Cerebras hardware.
Then I underline three potential bull cases:
The undiscovered TAM, where nobody knows what fast inference could be used for, so we are underpricing it today. I describe three potential market fits:
Fast fundamental investing
Embodied and human-like AI
Real-time human augmentation
The low-hanging fruit: architectural innovations that Cerebras can easily make, which would be a step function change in their unit economics. These include FP8/FP4 support and hybrid bonding.
Why they are the only choice for the non-NVIDIA ecosystem.
Most importantly, I culminate both yesterday’s and today’s research into a final conclusion and what I personally am doing at IPO.
Contents
5.3 Codex Spark
X Sentiment
My Own Testing: Equity Research
My Own Testing: Coding
Conclusion
Bull Case
The Undiscovered TAM
The Low Hanging Fruit
The Non-Nvidia Ecosystem
Conclusion and What I’m Doing at IPO
5.3 Codex Spark
5.3 Codex Spark is the custom-developed model by OpenAI, specifically made for running fast inference on Cerebras hardware. It is currently the only custom design model for Cerebras hardware, so in reality this model is one of, if not THE most important element of any Cerebras investment analysis. But for some reason, it is very under discussed.
It trades a large chunk of its performance to be smol. It features substantially lower benchmark performance, a context window that is one-eighth of the leading frontier models today, but with inference speeds of over 1,000 tokens per second.
As we established in the last post, in order to run a model of a certain size, a certain number of dinner plates must be strung together. The larger the model weights, the more wafer-scale engines are needed, and the more uneconomical it becomes. Therefore, small models are much more cost-efficient for Cerebras to run.
Now, how good is this model exactly? Here is where I found something funny.
The Cerebras IPO is one of the most hyped events of all time, but the only custom-designed model for its hardware, 5.3 Codex Spark, is one of the most un-hyped and un-loved models out there.
X Sentiment
AI lives on X, so let’s start with the X sentiment. Funny enough, I couldn’t find big accounts that posted about 5.3 codex spark just because of how underdiscussed it was, but there were enough posts for me to come to the conclusion that people generally did not like this model at all.
The main complaint is just the performance. It is a stupid model, to put it simply. Frontier intelligence unsurprisingly requires a certain amount of brute force that small models cannot muster.
My Own Testing: Equity Research
I have a little bit of a bias against benchmarks and testing that people usually do for these models. I don’t care about your score on Humanity’s last exam or whether you can build a snake game in one shot. These are not tasks that I will be doing in real life. This is a point people say a lot, but I think it is underappreciated because they are categorically different from benchmark tasks.
So, to that end, we will be doing two tests that I literally came up with by just asking myself what the next thing I have to do is. Which are equity research and coding.
In a very meta turn of events, I asked both 5.5 and 5.3 Codex Spark to look for pricing and usage data for 5.3 Codex Spark. So using Cerebras to evaluate Cerebras.
This is mainly a speed test. It is very difficult to gauge quality on research tasks because it is very subjective and could require large sample sizes all evaluated qualitatively before you can get a real sense of the difference. That is why we’re doing coding too.
But for speed, you might notice something interesting. Although Cerebras is supposed to be over ten times faster, 5.3 codex spark took 2 minutes and 19 seconds, while 5.5 took 3 minutes and 9 seconds. This means that 5.5 only took 36% more time rather than 900% more time.
This is because of two reasons. Both of which I observed first hand in my testing and is something us semiconductor people talk about all the time.
A large fraction, or maybe even a majority, of the time spent on doing the task was on web searches, tool calls, and JavaScript executions. Sound familiar? That’s right! CPUs strike again. All those research reports on how a large fraction of the latency in inference is CPU-bound. Well, turns out it isn’t just theory. 5.3 Codex does this really funny thing where it generates tokens in a millisecond, and then you have to wait two or three seconds for the CPU to run the tool call. You all should try it. It feels really weird, but you will immediately understand why CPUs are needed and why having fast inference isn’t all that it’s cracked up to be if you’re only speed boosting half of the workload.
Because the context window is so small, compaction happens very often. Every time the model has to compact its context, more latency is introduced. Even in that short two-minute task, GPT-5.3 Codex Spark had to compact its context, which is why it didn’t finish that far ahead of 5.5, which has a context window of a million tokens. Having a larger context window actually reduces latency by allowing the model to run without compaction for longer.
My Own Testing: Coding
I’m currently developing a full stack terminal for semiconductor investors, which I hope to release by the end of this month. It has a bunch of cool features like:
A Pokédex of every single AI infra company, split up by industry and niche
A map of all HPC co-location sites from the Bitcoin miners
A portfolio lab for backtesting and paper trading
An earnings calendar where you can view upcoming earnings, add events to your calendar, and download transcripts for past events
A Taiwan revenue tracker
A memory price tracker
An API endpoint for AI agents to access the app’s features autonomously
10 different calculators
(if you’re interested in beta testing for free subscription reach out)
Therefore, I assigned both models to build my next feature, which is a paper trading feature for the portfolio lab.
This was not a good test for speed because 5.5 ended up pulling up the in-app browser to test as if it were a real user, while 5.3 Spark just shipped it immediately. They made completely different workflow decisions, which means that it cannot be compared apple to apples, unlike for the equity research task. Even then, GPT 5.5 took 16 minutes, while 5.3 codex Spark took 3 minutes, which is only about a 5x difference. That’s still not 10x, even with one model doing way more!
How did they perform? Let’s look at 5.5 first. GPT 5.5 successfully implemented the paper trading feature, and I was able to create my new paper portfolio called Bloomentum.
5.3 Codex Spark, on the other hand, had this result.
I hope I do not have to explain any further why model size matters for real-world outcomes.
To see the actual finished product that 5.3 Codex Spark can create, I pasted the error back in and allowed it to fix the bug.
The quality of the work done by 5.3 Codex Spark is far worse. First of all there is another bug: the pricing quotes do not load at all. But if you just look at the UI of this tool, you can already notice the difference. It is very unintuitive and unfriendly to use. It is black and white instead of matching the colors of the app. The buy button isn’t even legible!
Additionally there’s this weird jittery effect where whenever you click a button or type something in, the elements sort of flicker.
Conclusion
Is 5.3 Codex Spark faster? Yes, without a doubt. However, is it 10 times faster like Cerebras claims? Not at all. Real-world workloads are not just pure feed forward. There are tool calls and there are context limitations. Just like doing well on the SAT doesn’t guarantee that you will be successful in a real-life career, running fast inference on a benchmark test in a lab does not guarantee that that same speedup applies in real life.
The performance difference, however, is extremely significant. There is a reason that people gravitate towards frontier-level intelligence even when cheaper open source alternatives exist. OpenAI and Anthropic’s combined ARR is not $80 billion for no reason. More intelligent models do things better. Not only do they do things better, they usually find ways to do them more efficiently, using less tokens in the process, and do it quicker as well. Not because of the sheer inference speed, but because of their superior reasoning abilities. The time it takes to correct bugs and failures, like the one seen in our test, will most likely overwhelm any gains from faster inference as the magnitude of the speed up in real life work just isn’t big enough.
And finally, one overlooked signal is the existence of 5.3 codex spark itself. OpenAI does not have an ultra-fast Cerebras option for GPT 5.5. They purposefully created a smaller weaker model optimized to run on this hardware. That’s actually pretty telling because it was a deliberate choice. They definitely thought about serving the frontier model on Cerebras first and then through testing and deliberation decided against it as it was either impossible or uneconomical.
Bull Case
Let’s be honest, these 5.3 codex spark tests were pretty bearish, so what’s the actual bull case? Well, I think I have a few.
The Undiscovered TAM
The undiscovered TAM argument simply posits that nobody really knows what fast inference could be used for, so the addressable market is completely undiscovered. Because nobody knows about these use cases, there is no market recognition of it, and because there’s no market recognition of it, it must be underpriced.
Here are three potential use cases that I can think of.
Fast Fundamental Investing
I’ve tried once or twice to trade earnings calls and catalysts and have failed every time. I come in with sheer, utter confidence in my fundamental understanding of the company and then get immediately humbled as my thinking speed is way too slow to process the information that was just presented to me and be able to connect that to what happens to the fundamentals of the company.
As you might imagine, this could be something that Cerebras is very good at. Imagine you’re a hedge fund and you have a lot of contacts in a company and a model already built. All you have to do is just give all of that to an LLM and tell it to watch out for catalysts. Once the event hits or the earnings call starts, the model is able to immediately run that inference and tweak the model and come to a conclusion. Having this fundamental-based conclusion just seconds or minutes before everyone else means that you can catch a 10 or 20% move while the rest of the market is still racking their brain around what just happened.
Embodied/Humanlike AI
What if I want an AI in my Zoom call? What if I want to talk to a model and have it interrupt me like my friends do? There is a level of speed required in humanlike back-and-forth conversation to enable idea exchange at a cadence that humans are familiar with.
Think about the way you talk to an AI right now. It is essentially long monologue after long monologue. This is due to the inference speed limitation. You may be able to imagine the plethora of use cases that can be unlocked with this constraint gone.
Robotics is another interesting application here. Imagine if humanoid robots can talk exactly like humans do. This would be an insanely large addressable market that includes all service industries like dining, hospitality, and front desk work.
Real-Time Human Augmentation
Finally, you have real-time human augmentation, which is essentially giving humans an earpiece that can allow an LLM to help that person in real time. Think about high-stakes meetings, live negotiations, that sort of thing.
The Low Hanging Fruit
The second category is low-hanging fruit. Low-hanging fruit being architectural innovations that Cerebras can easily make, which would be a step function change in their unit economics. Because their architecture is so much less mature than traditional GPUs, it is very unoptimized and therefore can be optimized.
FP8 & FP4 Support
The first thing that they can do is to enable FP8 and FP4 support. It is quite puzzling to me why they don’t do this already. They only currently support FP16, which is like having very limited storage on your phone, but the software only permits you to store photos and videos in 4K HD.
I think it’s likely because if they enabled this support, they would have to redesign their cores. Since their original design was meant for the training market, it was never optimized in this way. Anyways, there are rumors that this should be coming to the WSE-4, and that would unlock an easy 2x improvement in memory capacity.
Hybrid Bonding
The second thing they can do is more speculative. What if they hybrid bond an entire SRAM wafer on top of their WSE? It would be quite difficult, as it involves solving a ton of engineering problems and foundry shenanigans.
If they can do this, they scale their memory by an order of magnitude. Honestly, it would make Cerebras a completely different story.
But again, it is very speculative, and there’s no indication that they’re pursuing this.
The Non-Nvidia Ecosystem
Because Nvidia acquired Grok, they are the only major accelerated computing provider to have a fast inference solution. AMD and the hyperscalers are left without one. As you can see this argument is quite simple: the only way for them to compete effectively with Nvidia’s Grok LPU racks is to partner with Cerebras. All of Cerebras’s competitors are far behind and do not have comparable commercial traction, whether it’s MatX, SambaNova, or D-Matrix. Cerebras would be the only choice.
Conclusion and What I’m Doing at IPO
Because there are fixed, unchanging real-world constraints on total latency of agentic inference, there comes a phenomenon of diminishing returns on speeding up the feed forward itself, as there is no amount of pure transformer feed forward shenanigans that you can do that will make your CPUs do web searches, run JavaScript, and call tools faster. I believe Cerebras is well past that point. In my experience, GPT 5.5 low fast latency is already far more dependent on CPUs than the GPU.
The only model that is custom made for Cerebras is a subpar performer and has very little commercial traction. This makes me believe that the technical trade-off that they made and their “premium token” value proposition that we discussed in our last article does not currently have a market fit.
Therefore, I am staying out of the IPO tomorrow.
However, I will still be watching this company very closely because of the bull cases. None of them are currently unfolding today, but if they do, this entire story can flip on its head. The thing is, I am not going to speculate on them. I want to see them happen before I get in. If I miss the first 100% to guarantee that the thesis actually exists, that’s okay. Just look at the Bloom Energy or Lumentum charts. Anybody who missed the first 100% probably aren’t complaining about it.

















I struggle with this one. It seems like they can’t really bring up large models so if you take away this openai partnership, what models are they actually serving alone on this system. And how much standing capacity do you need for a model? Putting up large standing capacity for an open source model that turns over every three months is challenging. Fast tokens are sexy, but buyers need fungibility in the data center.
"I’ve tried once or twice to trade earnings calls and catalysts and have failed every time"
Biggest problem is that you can be right on the numbers and market still goes against you. Just infuriating when that happens lol