Skip to content

CLI reference

Flags always go before the model path; anything after it is the prompt or text. Every subcommand prints its own flags with -h.

Every size, on jitllm and jitllmd alike, is a byte count or a number with a unit: 4294967296, 4G, 512M.

Command What it does
jitllm hardware the CPU tier, every GPU each backend can open, and the spendable memory
jitllm library [-refs] the catalog convert fetches by name; -refs prints sizes and hf:// references for scripts
jitllm convert [-o DIR] [-q8] SOURCE [MMPROJ] OUT.jlm write a container from a GGUF, a Hugging Face directory or URL, or a catalog name. See Convert models.
jitllm run [flags] MODEL.jlm PROMPT... generate
jitllm embed [-lines] [-ids] MODEL.jlm TEXT... print an embedding model’s normalised vector as JSON
jitllm speed [flags] MODEL.jlm prompt and generation rates, warmed, as llama-bench measures them; or time to first token, or sessions decoded together
jitllm batch [flags] MODEL.jlm PROMPT... decode many sequences together on one device and report the aggregate rate
jitllm verify [flags] MODEL.jlm run on the CPU and a device and report where their logits differ
jitllm tokenize [-tokenizer FILE] MODEL TEXT... print token ids
jitllm info FILE.gguf a GGUF’s keys and tensors, before converting it
jitllm asm Q4_0|Q8_0|Q4_K[x4] disassemble the x86-64 kernel generated for a format on this machine
jitllm version the build’s version
Flag Default Meaning
-n N until the end tokens to generate; by default the model runs until it ends its reply or fills its context
-chat off apply the model’s chat template
-system TEXT a system message, with -chat
-image FILE an image for a vision model
-devices SPEC auto where to run; see the device grammar
-placement MAP where named blocks and the head run, with ! to pin and ~ to stream; see Placement maps
-placement-strict off fail when a -placement entry cannot be met, instead of placing that block elsewhere
-gpu-layers N -1 cap the blocks offered to devices; -1 fits as many as the card holds
-vram SIZE 0 device weight budget; 0 asks each device what is free and keeps an eighth of it, at least 256 MiB
-maxmem SIZE automatic host weight budget; see Memory and paging
-gpu-grow off start on the CPU and move blocks to the device while generating
-relocate on hand the device’s last block to the CPU when the context cannot grow on the device. On wherever -devices admits the host; -relocate=false turns it off
-tune-seam off measure whether fewer device blocks are faster, and move the seam
-no-gpu-fallback off fail instead of finishing on the CPU when the device breaks
-kv-cache DIR reuse KV pages of earlier prompts with a shared prefix, and spill history there under -kv-budget
-kv-cache-max SIZE 8 GiB cache size before least-recently-used eviction; 0 is unbounded
-kv-budget SIZE 0 KV history held in memory; older pages spill to -kv-cache and come back when read. 0 is unbounded
-depth N 0 pad the prompt to N tokens, to measure decode at that context length
-tokenizer FILE a tokenizer.json whose pre-tokenizer replaces the model’s
-spec mtp speculative decoding: draft with the model’s own multi-token-prediction block and verify the drafts in one pass; see Speculative decoding
-spec-k N 0 tokens drafted a round; 0 chooses each round from the measured acceptance and timing
-spec-min-p F 0 stop a round’s drafting at the first draft below this probability
-spec-rollback MODE auto how a hybrid’s recurrent state is taken back after a rejected draft: auto, rows or replay
-temp, -top-k, -top-p, -min-p, -repeat-penalty, -repeat-last-n, -seed greedy sampling, named as in llama.cpp; see Sampling
-cpuprofile FILE write a Go CPU profile
Flag Default Meaning
-devices SPEC auto as for run
-p N 512 prompt tokens
-n N 128 generated tokens
-r N 3 timed repetitions, after one untimed warm-up
-vram SIZE 0 device weight budget
-ttft off time to first token instead: cold (opening the model to the first token) and warm (a fresh session tokenizing, prefilling and sampling a -p-token prompt, -r times)
-sessions N decode N separate sessions -n tokens each, one after another and then together, interleaved over -r rounds
-gcstats off what the Go collector did during each timed phase: allocations a token, cycles, pauses, GC CPU
-cpuprofile FILE, -tgprofile FILE a CPU profile of the timed prompts, or of the timed decodes
-memprofile PREFIX allocation profiles around the first timed prompt and decode
-spec, -spec-k, -spec-min-p, -spec-rollback speculative decoding, as for run
Flag Default Meaning
-devices SPEC gpu the one device every block and the head go on
-batch N 16 sequences decoded together
-n N 128 tokens per sequence
-logits off read every row’s logits back and pick on the host, the path a sampled request takes, instead of the device’s argmax
-cpuprofile FILE write a Go CPU profile
Flag Default Meaning
-devices SPEC auto the device to compare with the CPU
-n N 32 tokens
-dlogit F 0 the largest per-logit difference to accept; an argmax flip inside it counts as a tie. 0 does not check
-migrate off move the CPU/GPU seam during the comparison
-prefill N 0 also compare the batched prefill of N tokens
-verify off recompute every served matrix product on the host and report disagreements
-vram SIZE, -maxmem SIZE automatic device and host budgets
-ab ARM run a paired A/B of one engine choice instead (split, submit, head, graph, softmax, ropetab, attnchunk, attnpair), with -rounds N and -depth N
Command Flag Meaning
convert -o DIR where a downloaded GGUF lands (default JITLLM_MODELS)
-q8 safetensors only: store every weight matrix as Q8_0
embed -lines embed each line of stdin, loading the model once
-ids print the token ids to stderr
library -refs name, size in bytes and hf:// references, one per line
tokenize -tokenizer FILE as for run
Command What it does
jitllmd serve [flags] run the server
jitllmd run [flags] PROMPT... stream a generation from a running server
jitllmd models list, load and unload models
jitllmd devices the server’s hardware and spendable memory
jitllmd sessions create, list, reset and close sessions
jitllmd place read and move a session’s blocks, tune its seam, and re-budget a model’s pager
jitllmd stats engine counters, once or streamed

Every client command takes -addr (host:port, :port or a full URL; default localhost:8080).

Flag Default Meaning
-addr :8080 listen address
-models DIR a developer path directory scanned by the model list and used to resolve a bare model name. Always pass it.
-load FILE a .jlm to load at startup, absolute or relative to -models
-id NAME derived the model name clients send
-devices SPEC CPU only devices for -load, in the device grammar
-maxmem SIZE engine’s own host page budget for -load, e.g. 8G
-gpu-layers N -1 at most N blocks on a device; -1 is as many as fit
-sessions N 1 concurrent sessions each device block reserves a KV cache for
-max-seq N model’s context default KV capacity per session, in positions
-host-concurrency N 1 sessions that may run on the CPU at once
-device-concurrency N 1 sessions that may run alone on one device at once; sessions wholly on a device batch instead
-max-batch N 0 requests of one device model that decode as rows of one step; 0 is the widest step the device runs, 1 turns batching off
-prompt-chunk N 0 prompt tokens a joining request feeds into one shared step; 0 is one device prefill chunk
-joint-steps MODE auto auto times a joint step against the rows one after another and runs the faster; always or never
-version S dev the version GetServerInfo reports
Flag Meaning
-model ID generate in an ephemeral session on this model
-session ID generate in an existing session, keeping its history
-continue continue from the session’s position instead of prefilling from scratch
-chat, -system TEXT apply the model’s chat template, with a system message
-n N tokens to generate (default: until the model ends its reply or fills the session’s context)
-stop A,B stop strings
-temp, -top-k, -top-p, -min-p, -repeat-penalty, -repeat-last-n, -seed sampling, as for jitllm run
-queue-timeout MS how long to wait for a device another session holds; 0 waits indefinitely
-echo have the server send the prompt’s tokens back first
-ids print the generated token ids to stderr
-quiet no started/finished lines on stderr
Flag Meaning
(none) list the model directory and what is loaded; -loaded lists only loaded models, -dir scans another directory
-load FILE load a container, with -id, -devices, -gpu-layers, -maxmem as for serve, and -kv-f16 to force an f16 KV cache
-unload ID unload a model; -force closes sessions holding it instead of refusing
Flag Meaning
(none) list sessions; -model filters
-new create a session on -model, with -id, -devices, -gpu-layers, -max-seq and -gpu-grow
-reset ID back to position 0
-close ID close it
-queue DEVICE the queue waiting on a device
Flag Meaning
-session ID report where each of the session’s blocks runs
-devices SPEC re-place the session’s blocks on these devices (they must be ones the model was loaded with); empty brings every block home
-gpu-layers N at most N blocks on a device
-head-on-device place the output projection on a device too
-move FIRST:LAST -to WHERE relocate a block range to host or a device
-relocate adopt one block per token as device memory frees
-tune-seam measure whether fewer device blocks are faster and move there, with -tokens, -rounds and -warmup
-model ID report the model’s pager residency; with -maxmem SIZE, re-budget it
Flag Meaning
-watch stream a snapshot every -interval milliseconds (default 1000)
-model ID, -session ID only this model or session

Read by jitllm only; jitllmd and the Go packages read no environment variables.

Variable Effect
JITLLM_MODELS default directory for downloaded models
HF_TOKEN token for private or gated Hugging Face repositories (HUGGING_FACE_HUB_TOKEN and ~/.cache/huggingface/token also work)
JITLLM_GPU_STREAM=1 place mixture blocks without their expert banks and upload only the chosen experts each token
JITLLM_ARENA=BYTES bound the host copy that streamed blocks are uploaded from; unset, a streaming placement sizes it
JITLLM_KV_F16=0 or 1 force the KV cache’s width instead of the per-CPU-tier default (f16 on AVX2, f32 elsewhere)
JITLLM_CORESET, JITLLM_CORES=N which cores decode uses (p by default) and how many of them
JITLLM_NUMA=off leave weight memory to the kernel’s first-touch policy instead of interleaving across NUMA nodes
JITLLM_HUGEPAGES=0 keep weight memory on the kernel’s default page size
JITLLM_OFFHEAP=0 put weight memory back on the Go heap (a control for measurements)
GOMAXPROCS when set, overrides the thread count jitllm aligns to its decode cores

The rest of the JITLLM_* names cmd/jitllm reads are measurement switches for engine development. The Go packages take options instead; see Embed it in Go.