Local-first AI runtime.
Built for inference
that doesn't wait.
Native transformer execution. ONNX model compilation. Distributed GPU cluster routing. Run any GGUF model locally with 12ms TTFT and zero third-party runtime overhead.
Built from first principles.
No third-party runtime wrappers. Deterministic memory. Laptop to GPU cluster.
Native Core
Go/C++ execution engine, GQA attention, Paged KV Cache 16-token blocks, SIMD CPU/CUDA streams.
EQC Compiler
ONNX graph compilation, operator fusion, dead node elimination, `.eqx` binary packaging.
EQ Cloud & Cluster
Multi-node GPU coordination, 15s heartbeat watchers, circuit breaker gateway, Kubernetes HPA.
Native Transformer Engine
Grouped-Query Attention (GQA) & Multi-Query Attention (MQA) implemented directly in Go/C++. Dynamic Paged KV Cache allocates fixed 16-token page blocks to avoid VRAM memory fragmentation.
func (n *NativeCore) ExecuteGQA(q, k, v *Tensor) *Tensor {
scale := 1.0 / math.Sqrt(float64(n.HeadDim))
kPages := n.KVCache.AppendKeyPages(k)
vPages := n.KVCache.AppendValuePages(v)
scores := MatMul(q, kPages.Transpose())
scores.Scale(scale)
return MatMul(Softmax(scores), vPages)
}EQC Model Compiler
Convert ONNX graphs into optimized .eqx binary packages.
[EQC] Ingesting ONNX model computation graph... [EQC] Pass 1: Constant Folding & Dead Node Elimination (14 nodes removed) [EQC] Pass 2: Fusing MatMul + SiLU -> SwiGLU Kernel [EQC] Pass 3: Memory Workspace Planning (Peak VRAM: 5.6 GB) [EQC] Serializing package model.eqx (Magic: 0x45515831)
Pre-evaluates static weights
SwiGLU & Attention kernel fusion
Workspace layout & stride planner
Binary package serialization
Six languages. One runtime.
Production-ready SDKs with full OpenAI API spec compatibility.
The full platform.
End-to-end tooling for local-first inference, model compilation, and cloud cluster serving.
EQ Engine
Runtime CoreLocal inference engine with CUDA/Metal acceleration, GQA attention, Paged KV Cache, and OpenAI REST API compatibility.
EQC Compiler
ONNX CompilerCompiles ONNX neural networks into standalone .eqx binary packages with fused SwiGLU kernels and planned memory layouts.
EQ Studio
Web ConsoleReal-time telemetry console monitoring VRAM residency, GPU utilization, streaming prompt completions, and DevTools.
EQ Cloud
Distributed ClusterMulti-node GPU cluster orchestrator with 15s heartbeat watchers, circuit breaker edge gateway, and Helm/K8s manifests.
Open-source, MIT licensed.
The entire EQ Engine runtime, EQC compiler, and SDK suite is open source. Fork it, extend it, ship it.
