← All posts / Tools

Kog's Bet: 30x Faster LLM Inference Without Buying a Single New GPU

French startup Kog hit 3,000 tokens per second on stock AMD MI300X and Nvidia H200 GPUs with pure software optimization — and now it's racing to bring that speed to full-size LLMs by September.

Kog's Bet: 30x Faster LLM Inference Without Buying a Single New GPU

The inference race has a new contrarian. While Cerebras counts its post-IPO winnings and a wave of startups design custom silicon to dethrone the GPU, an 11-person French startup called Kog is making the opposite bet: the GPUs already sitting in enterprise data centers have far more speed locked inside them than anyone has bothered to unlock — and software alone can set it free.

In an interview published August 14, 2026, Kog founder and CEO Gaël Delalleau laid out the company’s roadmap: prove that “30x faster LLM inference” is achievable on the AMD MI300X and Nvidia H200 GPUs that companies already own, then scale the technique from a 2-billion-parameter demo model to the full-size LLMs that businesses actually want to run.

3,000 tokens per second, on hardware you already have

Kog first turned heads in May 2026 when its tech preview hit the front page of Hacker News. The claim: extremely fast single-request decoding is possible on standard datacenter GPUs — no exotic hardware required. The demo hit 3,000 output tokens per second per single request on an AMD MI300X, a figure the company has also replicated on Nvidia’s H200. By comparison, typical frontier-model decoding runs in the tens of tokens per second for a single stream.

The catch was the model. The demo ran on Laneformer 2B, a purpose-built small model of roughly 2 billion parameters that Kog has since open-sourced. Skeptics quickly noted that a 2B model is a far cry from the 200B-plus frontier models where inference costs actually hurt. Delalleau’s answer: size is an engineering problem, not a physics problem.

Why the bet isn’t as crazy as it sounds

The received wisdom in the inference world is that GPUs are bottlenecked by memory bandwidth during decoding — which is exactly the argument used to justify custom silicon like Cerebras’ wafer-scale engines or the flood of new inference accelerator startups. Delalleau thinks that argument has quietly become a misconception.

Newer GPUs, he argues, ship with more and more memory bandwidth that “only begs to be unlocked.” The problem isn’t the hardware; it’s that mainstream inference stacks leave most of the hardware’s capability on the table. Kog’s Kog Inference Engine (KIE) digs down to a level below where most frameworks operate, treating each GPU architecture as a system to be reverse-engineered rather than a black box to be called through CUDA libraries.

Kog isn’t entirely alone on the software-first path. Fellow French startup ZML builds hardware-agnostic inference software that bypasses CUDA entirely. But Delalleau places Kog closer to Stanford’s Hazy Research lab, with an even deeper, hardware-level focus on GPU acceleration.

A hacker’s approach to silicon

The methodology comes directly from Delalleau’s unusual background. A solid-state physics graduate of France’s École Polytechnique, he spent years in offensive cybersecurity — white-hat hacking — and was a four-time finalist at DEFCON’s Capture the Flag competition. On the science side, he says, the mindset is “understanding the laws of physics, and the laws of the GPU in order to make the most of them.” Hacking taught him to reverse-engineer things “down to assembly language and binary code… to try to use it to achieve a goal for which it wasn’t necessarily designed.”

That depth has a cost. For every new GPU, Kog dedicates weeks or even months of GPU engineering research to that specific chip. With a team of 11, that caps how many architectures the company can support — at least until it can feed its methodology into agent-based pipelines that automate parts of the porting work.

The market is waiting

Demand signals have been strong. Kog pulled in 200 tangible business leads after its May preview, and Delalleau expects software engineering to be the first real beachhead. The pain is familiar to any Claude Code power user waiting hours for agent runs to finish — and Anthropic clearly believes speed is worth money, charging a price multiple for Claude’s Fast Mode. Design partners building prompt-to-game and prompt-to-app products also see direct revenue upside from faster generation.

One surprise from customer conversations: prospects aren’t prepared to fine-tune small models to fit accelerated inference profiles. That feedback pushed Kog to focus fully on accelerating large models instead of the small-model path of least resistance.

The company is backed by a seed round co-led by Varsity VC — the Paris early-stage fund founded by Delalleau’s former Stribe co-founder Kamel Zeroual — with support from Scaleway, France’s Bpifrance, and the French Tech 2030 program. As Europe pushes for sovereignty in both silicon and models, a French team extracting more from American-made GPUs fits the moment.

September is the moment of truth

Everything now hinges on one milestone: demonstrating the approach on a major LLM at 10x speed, which Delalleau targets for September 2026. That proof point unlocks both customer traction and, the company hopes, a Series A. Until then, Kog remains a fascinating test of a simple question the inference industry has mostly stopped asking: what if the fastest path to faster AI is not new chips, but understanding the ones we already have?