Models for Every AI Workflow
Run official EverestQ models or thousands of open models using one local-first runtime.
Official EverestQ Models

General-purpose flagship reasoning model with multi-step chain-of-thought execution and 128K context window.

Ultra-fast low-latency assistant model optimized for real-time conversational streaming and sub-10ms response times.

Mathematical proof and formal logic solver trained on scientific papers, theorem provers, and symbolic algebra.

High-throughput code generation, automated refactoring, and architectural design engine for complex software projects.

Lightweight edge model engineered specifically for mobile devices, laptops, and resource-constrained micro-servers.
Supported Open Models & Ecosystem (14)
Open-weight runtime executionGeneral-purpose flagship reasoning model with multi-step chain-of-thought execution and 128K context window.
Ultra-fast low-latency assistant model optimized for real-time conversational streaming and sub-10ms response times.
Mathematical proof and formal logic solver trained on scientific papers, theorem provers, and symbolic algebra.
High-throughput code generation, automated refactoring, and architectural design engine for complex software projects.
Lightweight edge model engineered specifically for mobile devices, laptops, and resource-constrained micro-servers.
Open-weights reasoning model with advanced chain-of-thought processing and state-of-the-art math/coding performance.
Meta flagship 70-billion parameter open model offering state-of-the-art general intelligence and instruction adherence.
Specialized code generation model trained on multi-trillion tokens of source code, documentation, and unit tests.
Google next-generation open weight model built from the same research as Gemini 2 models.
Microsoft state-of-the-art 14B parameter reasoning model trained with synthetic data curation techniques.
Mistral AI efficient 24B instruction-tuned model designed for fast low-latency enterprise workloads.
High-performance text embedding model for RAG semantic search and vector database indexing.
Multi-lingual, multi-functionality embedding model supporting dense, sparse, and multi-vector retrieval.
Ultra-small language model engineered for local browser simulation and low-footprint background services.