GPU Inference Speed Optimization: Kog Claims 30x Faster LLMs

5 Min Read

GPU inference speed optimization is at the heart of French startup Kog’s ambitious claim: 30x faster LLM inference on standard datacenter GPUs. The company’s tech preview demonstrated 3,000 tokens per second, and with 200 business leads already in hand, Kog is betting that software optimization can unlock more power from existing hardware.

The race for faster AI inference is intensifying, with custom hardware like Cerebras gaining attention. However, Kog is betting that significant GPU inference speed optimization can be squeezed from conventional GPUs already in enterprise data centers. By focusing on deep-level software optimization, the startup aims to unlock new capabilities on existing hardware.

Kog made waves in May with a tech preview that demonstrated “extremely fast single-request decoding” on standard datacenter GPUs like the AMD MI300X and Nvidia H200. While this didn’t include laptop GPUs, the potential for unlocking more power from existing hardware attracted over 200 tangible business leads.

The Science Behind GPU Inference Speed Optimization

Kog’s initial demo showed an impressive 3,000 per-request tokens per second (TPS), but this was achieved using a purpose-built small model, Laneformer 2B. The company’s ultimate goal is to deliver “30x faster LLM inference” on much larger models.

CEO Gaël Delalleau acknowledges the skepticism but remains confident. “GPUs have a bright future,” he told TechCrunch. He argues that the idea of GPUs being poorly suited for decoding is a misconception, as newer chips offer more memory bandwidth that GPU inference speed optimization techniques can unlock.

Software Engineering as First Target

Based on early feedback, Kog expects software engineering to be its first major use case. Many developers using tools like Claude Code are frustrated by long wait times for results. Kog’s Inference Engine (KIE) aims to dramatically reduce these delays through GPU inference speed optimization, which is critical for users relying on AI for professional workflows.

Kog is also working with design partners in fields like game and app development, where faster inference times could lead to more revenue. The company is adapting its approach based on market demand, focusing on accelerating larger models rather than requiring customers to fine-tune smaller ones.

A Deep-Level Technical Approach

Kog’s methodology is deeply technical, rooted in its founder’s unique background. Delalleau studied solid-state physics at France’s École Polytechnique and worked in offensive cybersecurity. This combination of understanding the laws of physics and reverse-engineering at the assembly level shapes the company’s approach to GPU inference speed optimization.

Learning from the Hardware

The Kog team dedicates weeks or even months to deeply understand each new GPU architecture. “For every new GPU, we’ll dedicate several weeks or even months to really dig into the details and conduct GPU engineering research on that hardware,” Delalleau explained. With a team of 11, this limits the number of chips Kog can support for GPU inference speed optimization in the short term.

Kog isn’t alone in this space. ZML, another French startup, offers hardware-agnostic software that bypasses Nvidia’s CUDA for fast inference. Delalleau compares Kog more closely to Stanford’s Hazy Research, emphasizing their even deeper-level focus on acceleration.

Kog is supported by Scaleway, France’s Bpifrance, and the French Tech 2030 program, which adds potential sovereignty tailwinds as Europe builds its own AI capabilities.

The company’s immediate goal is to prove its GPU inference speed optimization approach works on large language models. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau stated.

Kog’s long-term vision includes feeding its methodology into agent-based pipelines to support more chips and models. For now, the startup must demonstrate that its unique, hands-on approach to GPU inference speed optimization can deliver results on the scale the market demands.

Share This Article
Leave a Comment