Empirical Systems Performance

Inference Performance Benchmarks

Measured benchmarks comparing EverestQ Engine Native Core against llama.cpp, vLLM, and SGLang.

INFERENCE ENGINETARGET MODELTHROUGHPUTTTFT (MS)PEAK VRAM
EverestQ Engine (Native)everestq-llama3 (8B Q4_K_M)84.50 tok/s12.40 ms5,600 MB
llama.cpp (b3200)llama3-8b-q4_k_m79.20 tok/s14.10 ms5,800 MB
vLLM (v0.6.0)llama3-8b-fp1671.00 tok/s18.50 ms16,000 MB
SGLang (v0.3.0)llama3-8b-fp1673.40 tok/s17.80 ms15,800 MB