EQ Engine Logo
EQ Enginev1.0
Back to Model Library
EverestQ

Kumari

Gold Verified
Publisher: EverestQ Engineering · Category: Chat
Parameters
8B
Context Window
32K
Default Size
5.6 GB
Downloads
289,100

Ultra-fast low-latency assistant model optimized for real-time conversational streaming and sub-10ms response times.

PULL MODEL
eq pull everestq/kumari
RUN INFERENCE
eq run everestq/kumari "Summarize quantum computing principles"
Attention Architecture

Utilizes Grouped-Query Attention (GQA) with 16-token fixed Paged KV Cache block allocation for zero-fragmentation memory residency.

ONNX & EQC Compatible

Compiles directly through the EQC toolchain into standalone .eqx binary packages with fused SwiGLU kernels.

Multi-Hardware Support

Auto-detects CUDA RTX/A100/H100, Apple Silicon Metal Performance Shaders, or AVX-512 CPU execution backends.