Skip to content

How it works

Most inference engines are a set of hand-written kernels and a loop that calls them. jitllm is closer to an operating system. It treats a model the way an OS treats a process, and the machine’s CPUs, GPUs and memory as resources to schedule it onto.

An operating system jitllm
a process a loaded model
a thread a block of the model
CPU cores CPU cores and every accelerator
the scheduler which block runs where
a physical page a block, an expert or a run of KV positions, in memory
swap the .jlm file on disk
a page fault a page-in

It is one process: the analogy is about what it decides, not how it is built.

Convert once, compile every time. The container fixes the layout; the kernels are generated at load for whatever CPU and GPUs the runtime finds.

There is no library of precompiled kernels. At load, jitllm probes the CPU and every GPU, then emits machine code for the model’s actual shapes and weight formats: x86-64 or ARM64 instructions for the host, PTX, SPIR-V or Metal for the devices. The assemblers are part of jitllm, written in Go.

Specializing per machine and per model is what makes a kernel fast here. Shapes are constants, so loops unroll to exact widths; register allocation is planned for the model’s dimensions; and the code uses the instructions this CPU has (VNNI, or a sequence that does the same job on a CPU without it).

Then it measures. Tuners time real kernels on this machine and cache what wins: how many rows each worker takes, how many tokens a prefill tile holds, how a matrix product splits across GPU threads, how many cores a region wants. A choice that depends on the machine is measured on the machine, not read from a table.

target
formatk
  1. probe
  2. emit
  3. map
  4. tune
  5. run
sketch of the emitted inner loop
tuner

from the hardware probefrom the weight formata constant from the shape
The code is written for this machine and this model, at load. Change the target, format or width and the loop is emitted again: other instructions, other unpacking, another trip count baked in as a constant. This is a sketch of the inner loop's structure, not a disassembly; jitllm asm Q4_K prints the real x86-64 kernel.

Several engines now say they compile on your machine. They do not compile the same thing, and the difference is who writes the kernel.

jitllm Triton (vLLM, SGLang) Magnitude
Who writes a kernel jitllm’s emitters, Go code that writes the kernel for the shape, format and instructions in front of it A person, in Triton’s Python language, one function per kernel A person: “hand-optimized kernels for popular open-weight families”, as templated source
What turns it into machine code jitllm’s own assemblers: x86-64 and ARM64 bytes directly; PTX, SPIR-V and MSL that the GPU driver finishes Triton’s compiler, through MLIR and LLVM, then the GPU vendor’s assembler NVRTC, NVIDIA’s run-time C++ compiler, which it ships; the operating system’s Metal compiler
When At load, for this model on this machine The first call with new specializations, then cached on disk On the device, before a model runs
What is baked in The model’s shapes as constants, its weight formats, the instruction set this CPU or GPU has The arguments the author marked constant, types, alignment The parameters the author declared
The CPU Generated at run time, like the GPU Not covered: Triton writes GPU kernels “CPU kernels are compiled into the binary”
Tuning Times variants on real work on this machine and keeps the fastest Times the configurations the author listed (@triton.autotune) “A time-bounded local search of the author’s declared parameter domain”
Needed at run time Nothing but the GPU driver Python, PyTorch, and Triton with its LLVM Its bundled NVRTC

Triton and Magnitude compile kernels a person wrote, at run time, with a compiler. jitllm generates the kernel: a person writes the emitter, and the emitter writes the code for whatever model and hardware it meets, on the CPU as well as the GPU, with no compiler in the loop on the CPU and no kernel source to fill in. So a new shape needs no new code at all, a new weight format or instruction set needs one emitter case rather than a kernel per model family, and a model nobody has tuned for still runs entirely as generated code. torch.compile comes closest on the Python side: it generates Triton kernels from a PyTorch graph, which Triton then compiles. Magnitude’s description is from its README and its compilation design; Triton’s from its documentation.

There is no interpreted fallback. A shape without an optimized kernel gets a basic generated one, on the CPU or the device, so a model never quietly drops into slow Go loops.

A transformer is a stack of repeating blocks, and a block is the unit everything below is decided in. Conversion writes each block as one fixed-size page of the .jlm file, already in the layout every backend reads, so bringing a block into memory is a read and nothing else.

Opening a model reads no weights. The first token reads each block as it reaches it; a budget that holds the whole model never evicts, so from then on a token reads nothing from disk.

Opening reads nothing; the first token reads each block once. With a budget that holds the model, every later token reads nothing from disk.

jitllm keeps the graph (embedding, sampling, the order of blocks) on the host and hands devices whole blocks. A device runs a block’s attention, norms, feed-forward and expert routing without returning to the host between them, so a token crosses between the card and the host once, at the seam.

The first blocks go to the fastest device, the rest to the next, and whatever no device takes runs on the CPU. A block on the card has no second copy in host RAM. See Devices and placement for choosing devices and writing a placement map.

Drag the seam. A token runs the card's blocks, crosses the seam once, and finishes on the host. A block on the card has no copy in host RAM.

Every token leaves a row of attention history in every block. jitllm stores it in fixed-size pages that are allocated as the context grows, not reserved for the longest context up front, and each block’s pages live on whichever side the block runs: on the card in a per-layer page pool, on the host in RAM.

block weightsa KV page (4 positions)the tokenread from disk
History is allocated a page at a time. Each block keeps its own KV pages, on the side the block runs. A short conversation holds a few pages, not a reservation for the whole context.

When host RAM holds fewer block pages than the host runs blocks, the missing ones are read from the file as the token reaches them. Block pages are evicted most recently used first, which is optimal for a model scanned front to back every token: one page short of fitting costs one page-in a token. Least recently used, the textbook default, would evict exactly the block needed next and read the whole host side every token.

block weightsa KV page (4 positions)the tokenread from disk
Set host RAM one page short. Under MRU a token pays one page-in; under LRU every fault evicts the block the token needs next, and the whole host side is read every token.

A mixture of experts pages at a finer grain: each expert is its own page, evicted least recently used, so a token reads only the experts its router chose and does not already hold.

one mixture block: 32 experts, the router picks 4 a token
this token read 0 expert pagesaverage 0 a tokena whole-block page would read 32
resident expert pagerouted this tokenread from disk this token
A token reads what it routes to, and only if it is not already there. On Moonlight-16B-A3B under a 6 GB budget, giving each expert its own page took CPU decode from 0.890 to 0.144 GiB read a token, and ran 3.1× faster.

KV pages spill when history outgrows memory

Section titled “KV pages spill when history outgrows memory”

-kv-budget caps how much history a session holds in memory. Older KV pages move to the -kv-cache directory, and since attention reads the whole history every token, they stream back as it reads them. A context longer than memory runs, more slowly, instead of failing. The same directory keeps prompt prefixes for reuse; see Memory and paging.

block weightsa KV page (4 positions)the tokenread from disk
Lower -kv-budget. The oldest pages move to the -kv-cache directory; attention still reads all of the history, so they stream back each token, outlined as they are read.

The seam can move between tokens, and a block’s KV pages move with it, so nothing is recomputed. -relocate uses that when the card fills: as the context grows, the card’s history takes the room its weights had, and the last block on the card moves to the host instead of the request failing. The host then pages what it now holds, by the same rules as before.

block weightsa KV page (4 positions)the tokenread from disk
Let the conversation grow. When the card's weights and history no longer fit, -relocate hands its last block to the host, KV pages and all. Turn it off and the context stops growing on the card.

A server holds several models at once and many sessions per model. Each session owns its sequence (its KV cache, its position and where its blocks run) while the model’s weights are shared. On a device, each session’s history lives in its own pages of a shared pool, so sessions never see each other’s state. Requests on a device decode together, as rows of one step that reads each block’s weights once for all of them, and a joining request’s prompt rides in the same steps a chunk at a time, so it never stalls the rows already generating.

Every generated kernel is tested against a reference implementation, including ragged shapes and a deliberate violation to prove the test can fail. Each architecture is checked against an outside reference, llama.cpp or Hugging Face transformers, on real checkpoints: token by token, and on the logits themselves where a greedy transcript could hide a bug. Device placement is checked end to end against the host on models that have the feature being placed.

The repository records the measurements behind each design decision, including the ones that were tried and did not pay: