EQ Engine Logo
EQ Enginev1.0

Local-first AI runtime.
Built for inference
that doesn't wait.

Native transformer execution. ONNX model compilation. Distributed GPU cluster routing. Run any GGUF model locally with 12ms TTFT and zero third-party runtime overhead.

12.40 ms TTFT·84.50 tok/s throughput·<45 ms cold start
eq serve — bash
localhost:11434
❯
GGUF Binary◆ONNX Graphs◆CUDA Acceleration◆Metal (Apple Silicon)◆Go SDK◆TypeScript SDK◆Python SDK◆Rust SDK◆Docker Container◆Kubernetes HPA◆OpenAI Compatible◆Paged KV Cache◆.eqx Package Serializer◆GGUF Binary◆ONNX Graphs◆CUDA Acceleration◆Metal (Apple Silicon)◆Go SDK◆TypeScript SDK◆Python SDK◆Rust SDK◆Docker Container◆Kubernetes HPA◆OpenAI Compatible◆Paged KV Cache◆.eqx Package Serializer◆

Built from first principles.

No third-party runtime wrappers. Deterministic memory. Laptop to GPU cluster.

Native Core

Go/C++ execution engine, GQA attention, Paged KV Cache 16-token blocks, SIMD CPU/CUDA streams.

EQC Compiler

ONNX graph compilation, operator fusion, dead node elimination, `.eqx` binary packaging.

EQ Cloud & Cluster

Multi-node GPU coordination, 15s heartbeat watchers, circuit breaker gateway, Kubernetes HPA.

12.40 ms
Time to First Token
84.50
Throughput (tok/s)
<45 ms
Cold Daemon Start
5600 MB
Peak VRAM (CUDA)

Native Transformer Engine

Grouped-Query Attention (GQA) & Multi-Query Attention (MQA) implemented directly in Go/C++. Dynamic Paged KV Cache allocates fixed 16-token page blocks to avoid VRAM memory fragmentation.

◆Quantization formats: FP32, FP16, BF16, INT8, INT4, Q4_K_M
◆SwiGLU MLP activation kernels with vector SIMD instructions
◆Zero runtime overhead: laptop CPU, CUDA GPU, or Apple Metal
internal/native/attention.goGo
func (n *NativeCore) ExecuteGQA(q, k, v *Tensor) *Tensor {
    scale := 1.0 / math.Sqrt(float64(n.HeadDim))
    kPages := n.KVCache.AppendKeyPages(k)
    vPages := n.KVCache.AppendValuePages(v)
    scores := MatMul(q, kPages.Transpose())
    scores.Scale(scale)
    return MatMul(Softmax(scores), vPages)
}

EQC Model Compiler

Convert ONNX graphs into optimized .eqx binary packages.

eqc compile pipelineTerminal
[EQC] Ingesting ONNX model computation graph...
[EQC] Pass 1: Constant Folding & Dead Node Elimination (14 nodes removed)
[EQC] Pass 2: Fusing MatMul + SiLU -> SwiGLU Kernel
[EQC] Pass 3: Memory Workspace Planning (Peak VRAM: 5.6 GB)
[EQC] Serializing package model.eqx (Magic: 0x45515831)
1Constant Folding

Pre-evaluates static weights

2Operator Fusion

SwiGLU & Attention kernel fusion

3Memory Planning

Workspace layout & stride planner

4.eqx Output

Binary package serialization

Six languages. One runtime.

Production-ready SDKs with full OpenAI API spec compatibility.

Go
sdk.NewClient()
Python
everestq-python
TypeScript
@everestq/sdk
Rust
everestq-rs
Java
ai.everestq
C#
EverestQ.SDK

The full platform.

End-to-end tooling for local-first inference, model compilation, and cloud cluster serving.

EQ Engine

Runtime Core

Local inference engine with CUDA/Metal acceleration, GQA attention, Paged KV Cache, and OpenAI REST API compatibility.

EQC Compiler

ONNX Compiler

Compiles ONNX neural networks into standalone .eqx binary packages with fused SwiGLU kernels and planned memory layouts.

EQ Studio

Web Console

Real-time telemetry console monitoring VRAM residency, GPU utilization, streaming prompt completions, and DevTools.

EQ Cloud

Distributed Cluster

Multi-node GPU cluster orchestrator with 15s heartbeat watchers, circuit breaker edge gateway, and Helm/K8s manifests.

Open-source, MIT licensed.

The entire EQ Engine runtime, EQC compiler, and SDK suite is open source. Fork it, extend it, ship it.