Mistral Small
Partner VerifiedPublisher: Mistral AI · Category: Chat
Parameters
24B
Context Window
32K
Default Size
14.8 GB
Downloads
215,900
Mistral AI efficient 24B instruction-tuned model designed for fast low-latency enterprise workloads.
PULL MODEL
eq pull mistral-small
RUN INFERENCE
eq run mistral-small "Synthesize meeting notes into action items"
Attention Architecture
Utilizes Grouped-Query Attention (GQA) with 16-token fixed Paged KV Cache block allocation for zero-fragmentation memory residency.
ONNX & EQC Compatible
Compiles directly through the EQC toolchain into standalone .eqx binary packages with fused SwiGLU kernels.
Multi-Hardware Support
Auto-detects CUDA RTX/A100/H100, Apple Silicon Metal Performance Shaders, or AVX-512 CPU execution backends.
