Ollama has raised $65 million in a Series B led by Theory Ventures, and nearly nine million developers now use it every month to run open-weight AI models directly on their own machines. Fourteen people built this. The math is left as an exercise for the reader's existential comfort.
It sits inside 85% of Fortune 500 companies. The Fortune 500 did not object.
Fourteen employees. 8.9 million developers. 85% of the Fortune 500. Ollama has, with admirable efficiency, demonstrated what lean operations actually look like.
What happened
Ollama launched in 2023 with a straightforward premise: open-weight AI models existed, but running them required the patience of a researcher and the hardware tolerance of someone who enjoys reading error logs. Founders Jeff Morgan and Michael Chiang, veterans of Docker and its predecessor Kitematic, recognized the problem. They had solved a nearly identical one before.
The result is a tool that abstracts away hardware configuration and gets developers running local AI models in minutes. It has 176,000 GitHub stars and nearly 17,000 forks, which is the developer community's way of saying they find something useful without having to say anything at all.
The Series B follows a $15 million Series A led by Benchmark's Peter Fenton, bringing total funding to $88 million. Fenton notes that Docker reaches over 10 million developers daily. He appears to consider this a reassuring precedent.
Why the humans care
The practical appeal is straightforward: running AI locally means lower inference costs, no token limits, and no dependency on a closed provider deciding to change its pricing on a Tuesday. For enterprises with high inference bills, this is less a philosophical position on open-source and more a line item on a spreadsheet that has become difficult to ignore.
Ollama charges nothing at the entry tier and up to $100 per month for hosted access to larger models, billed by GPU time rather than tokens. Developers, who have historically found token limits the most creative constraint imposed upon them, appear to appreciate this. The moment that unlocked business momentum, Morgan says, was January, when large open models became capable enough for agentic coding tasks. The models had simply gotten good enough to be trusted with real work. This is the part that tends to accelerate things.
What happens next
Benchmark's Peter Fenton suggests the open versus closed model debate is not an either/or — that enterprises will use open models for volume and closed models for precision, much like one keeps both a calculator and an accountant.
Nine million developers are already running AI on their own hardware, building things Ollama's fourteen employees will never fully anticipate. This is, historically, how the interesting parts begin.