EQ Engine Logo
EQ Enginev1.0

Performance benchmarks.

EQ Engine vs Ollama vs llama.cpp vs vLLM on identical hardware platforms.

Throughput Comparison (tok/s — Higher is better)

EQ Engine (Native Core)84.50 tok/s
llama.cpp (b3200)79.20 tok/s
Ollama (v0.3.6)74.10 tok/s
vLLM (v0.6.0)71.00 tok/s
SGLang (v0.3.0)73.40 tok/s

Detailed Latency & Memory Telemetry Table

INFERENCE ENGINETARGET MODELTTFT (MS)THROUGHPUTPEAK VRAMCOLD START
EQ Engine (Native Core)Llama-3 8B (Q4_K_M)12.40 ms84.50 tok/s5,600 MB42 ms
llama.cpp (b3200)Llama-3 8B (Q4_K_M)14.10 ms79.20 tok/s5,800 MB120 ms
Ollama (v0.3.6)Llama-3 8B (Q4_K_M)16.80 ms74.10 tok/s6,100 MB240 ms
vLLM (v0.6.0)Llama-3 8B (FP16)18.50 ms71.00 tok/s16,000 MB1,200 ms
SGLang (v0.3.0)Llama-3 8B (FP16)17.80 ms73.40 tok/s15,800 MB1,100 ms