# OpenLLMStack > The open source LLM ecosystem hub: models, inference engines, optimizations, and agentic frameworks. Every model page carries verified specs, published benchmark scores, current API pricing, and sourced practitioner commentary. Figures on this site are attributed to a primary source (a model card, technical report, provider pricing page, or independent evaluation) and dated, so they can be quoted directly. Prices are per 1 million tokens in USD unless stated otherwise. ## Models - [GLM-5.3](https://openllmstack.com/models/glm-5-3/): 753B (40B active) params, 1M context, GLM-5.3 License. GLM-5.3 is a 753B MoE from Z.ai with 1M context, 84.5% on CyberGym, and open FP8 weights at $4.40 per 1M output tokens. Specs, benchmarks, and pricing. - [DeepSeek-V4-Pro-0813](https://openllmstack.com/models/deepseek-v4-pro-0813/): 1.6T (49B active) params, 1M context, MIT license. DeepSeek-V4-Pro-0813 is an open 1.6T MoE with 1M context, 87.9 on Terminal Bench 2.1, and output at $1.98 off-peak per 1M tokens. Specs, benchmarks, and pricing. - [Qwen3.8-2.4T-A95B](https://openllmstack.com/models/qwen3-8-max/): 2.4T (95B active) params, 1M (262K native) context, Qwen3.8-Max License. Qwen3.8-2.4T-A95B is the open-weight Qwen3.8-Max: a 2.4T-param MoE with 262K native context, 86.6 on Terminal Bench 2.1, and API output at $6 per 1M tokens. - [Kimi-K3](https://openllmstack.com/models/kimi-k3/): 2.8T (104B active) params, 1M context, Kimi K3 License. Kimi K3 from Moonshot AI is an open 2.8T-param MoE with 1M context and native vision. It scores 88.3% on Terminal Bench 2.1 and ranks #1 on Frontend Code Arena. - [DeepSeek-V4-Flash-0731](https://openllmstack.com/models/deepseek-v4-flash-0731/): 284B (13B active) params, 1M context, MIT license. DeepSeek-V4-Flash-0731 is an open 284B MoE with 13B active params, 82.7 on Terminal Bench 2.1, and output at $0.28 per 1M tokens. Specs, benchmarks, and pricing. - [DeepSeek-V4-Pro](https://openllmstack.com/models/deepseek-v4-pro/): 1.6T (49B active) params, 1M context, MIT license. DeepSeek-V4-Pro is an open 1.6T-param MoE with 1M context, 80.6% on SWE-bench Verified, and output at $0.87 per 1M tokens. Full specs, benchmarks, and pricing. ## Guides and analysis - [The Best Open Source LLMs in 2026](https://openllmstack.com/blog/best-open-source-llms/): Compare the best open source LLMs in 2026 on parameters, context, license, price, and more. Output costs range from $0.87 to $15.00 per 1M tokens. Published August 10, 2026. - [Is DeepSeek Better Than ChatGPT? 25+ DeepSeek Stats (2026)](https://openllmstack.com/blog/deepseek-statistics/): Explore 25+ DeepSeek statistics for 2026, including downloads, users, usage, training cost, API pricing, funding, and ChatGPT comparisons. Published July 16, 2026. - [What Is KV Cache? A Plain-English Guide for LLM Inference](https://openllmstack.com/blog/what-is-kv-cache/): Learn how KV caching speeds up LLM inference, how to calculate memory use, and how to reduce GPU cost with quantization and offloading. Published July 13, 2026. - [50+ Open Source LLM Statistics & Trends (2026)](https://openllmstack.com/blog/open-source-llm-statistics/): Explore 50+ open source LLM statistics for 2026, covering ecosystem growth, model usage, performance, inference costs, and leading developers. Published July 6, 2026. - [About OpenLLMStack](https://openllmstack.com/blog/about-openllmstack/): Why I built OpenLLMStack and what you can expect: a single, curated hub for the open source LLM ecosystem. Published June 24, 2026. ## Timeline of open LLM milestones Full timeline: https://openllmstack.com/timeline/ - July 31, 2026, DeepSeek-V4-Flash-0731 Redraws the Cost Curve: DeepSeek retrains V4-Flash for agents and coding and ships it at the same price: $0.14 per million input tokens and $0.28 per million output tokens. Terminal Bench 2.1 jumps to 82.7 with 13B active parameters, putting an MIT-licensed model on the Pareto frontier at roughly 1% of the output price of the closed frontier. https://openllmstack.com/models/deepseek-v4-flash-0731/ - July 24, 2026, Tech Industry Open Weight Letter: After Kimi K3 hits public download and stirs concern that policymakers may crack down on open AI systems, Jensen Huang joins Satya Nadella and dozens of tech companies on a letter warning against restricting open models. https://x.com/JensenHuang/status/2080643682408321103 - July 16, 2026, Kimi K3 Released: Moonshot AI releases Kimi K3, the first open model to reach 2.8 trillion parameters. The native multimodal MoE model combines a 1M-token context window with Kimi Delta Attention and Attention Residuals. https://openllmstack.com/models/kimi-k3/ - June 16, 2026, GLM-5.2 Released: Zhipu AI releases GLM-5.2, a large MoE model under the permissive MIT license with a 1M context window, designed for long-horizon agentic tasks. It beats GPT-5.5 on several benchmarks and trail Claude Opus 4.8 slightly, at far lower cost. https://z.ai/blog/glm-5.2 - April 24, 2026, DeepSeek-V4 Released: DeepSeek releases V4-Pro and V4-Flash, trained on over 32T tokens and natively supporting a 1M-token context window. Built on Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to extend context and efficiency. https://api-docs.deepseek.com/news/news260424 - September 29, 2025, DeepSeek Sparse Attention: DeepSeek introduces Sparse Attention (DSA) with DeepSeek-V3.2-Exp, selectively attending to the most relevant tokens to cut the cost of long-context inference while preserving quality. https://arxiv.org/abs/2512.02556 - April 5, 2025, Llama 4 Released: Meta releases Llama 4 (Scout and Maverick), its first natively multimodal, mixture-of-experts model family with very long context windows. Pushes open models further into MoE architectures. https://ai.meta.com/blog/llama-4-multimodal-intelligence/ - January 20, 2025, DeepSeek-R1 — Open Reasoning: DeepSeek releases R1, the first open-weight reasoning model that rivals OpenAI's o1. Achieves chain-of-thought reasoning through pure RL without supervised reasoning traces. https://github.com/deepseek-ai/DeepSeek-R1 - December 26, 2024, DeepSeek-V3 Released: DeepSeek releases V3, a 671B MoE model trained for only $5.5M — shattering cost assumptions. Introduces Multi-head Latent Attention (MLA) and matches top proprietary models. https://github.com/deepseek-ai/DeepSeek-V3 - October 15, 2024, Agentic Frameworks Surge: 2024 sees an explosion of agentic frameworks — LangGraph, CrewAI, AutoGen, and others mature rapidly. The focus shifts from simple chatbots to autonomous multi-step agents. - July 23, 2024, Llama 3.1 405B Released: Meta releases Llama 3.1 including a 405B parameter model — the largest open-weight model at the time, competitive with GPT-4 class models on many benchmarks. https://ai.meta.com/blog/meta-llama-3-1/ - January 17, 2024, SGLang Launched: The LMSYS team releases SGLang, a fast serving framework introducing RadixAttention for automatic KV cache reuse across requests. Becomes a leading high-performance inference engine alongside vLLM. https://github.com/sgl-project/sglang - September 27, 2023, Mistral 7B Released: Mistral AI drops Mistral 7B via a torrent link with no fanfare. Outperforms Llama 2 13B on all benchmarks, introducing sliding window attention and grouped-query attention. https://mistral.ai/news/announcing-mistral-7b/ - July 18, 2023, Llama 2 Released: Meta releases Llama 2 with a permissive commercial license, including chat-tuned variants. First truly open model family that could be freely deployed in commercial products. https://ai.meta.com/llama/ - June 20, 2023, vLLM Introduces PagedAttention: UC Berkeley releases vLLM with PagedAttention, applying virtual memory concepts to KV cache. Achieves 2-4x throughput improvement and becomes the standard for LLM serving. https://blog.vllm.ai/2023/06/20/vllm.html - March 11, 2023, llama.cpp Launched: Georgi Gerganov releases llama.cpp, enabling LLaMA inference in C/C++ on consumer hardware including Apple Silicon. Democratizes local LLM access. https://github.com/ggerganov/llama.cpp - February 24, 2023, LLaMA Released by Meta: Meta releases LLaMA (7B-65B), proving smaller models trained on more data can compete with much larger ones. Catalyzes the open source LLM movement. https://arxiv.org/abs/2302.13971 - January 23, 2020, Scaling Laws for Neural Language Models: OpenAI publishes 'Scaling Laws for Neural Language Models', showing that model performance improves as a predictable power law with compute, data, and parameters — the empirical foundation for the era of scaling LLMs. https://arxiv.org/abs/2001.08361 - December 3, 2019, PyTorch Paper Published: The PyTorch team publishes 'PyTorch: An Imperative Style, High-Performance Deep Learning Library', detailing its define-by-run design. PyTorch becomes the dominant framework for LLM research and training. https://arxiv.org/abs/1912.01703 - February 14, 2019, GPT-2 Released: OpenAI unveils GPT-2, a 1.5B parameter language model whose surprisingly fluent text generation sparks debate over staged releases and the risks of large language models. https://openai.com/index/better-language-models/ - June 12, 2017, Transformer Architecture Published: Google Brain publishes 'Attention Is All You Need', introducing the Transformer — the architecture that would power all modern LLMs. https://arxiv.org/abs/1706.03762 ## Inference engines Overview: https://openllmstack.com/inference/ - [Atomic Chat](https://atomic.chat): Local AI app and inference engine for agents, running open-weight LLMs fully offline on desktop and mobile. Bundles llama.cpp, a TurboQuant KV-cache fork, and MLX-VLM for Apple Silicon behind one OpenAI-compatible server on localhost:1337, with MTP and DFlash speculative decoding and MCP tool support. - [llama.cpp](https://github.com/ggerganov/llama.cpp): C/C++ inference of LLMs with minimal dependencies. Supports GGUF quantization, runs on CPU and Apple Silicon with optional GPU offloading. The backbone of local LLM inference. - [LMDeploy](https://lmdeploy.readthedocs.io): Toolkit for compressing, deploying, and serving LLMs from the InternLM team. Its TurboMind engine delivers high request throughput with persistent batching, blocked KV cache, and weight-only quantization. - [MAX](https://www.modular.com/max): Modular's high-performance inference framework and serving platform. Pairs a Mojo-powered kernel stack with an OpenAI-compatible server for fast, portable LLM serving across NVIDIA and AMD GPUs. - [MLC LLM](https://llm.mlc.ai): Universal LLM deployment engine built on Apache TVM. Compiles and runs models natively across GPUs, browsers (WebGPU), iOS, and Android for high-performance inference everywhere. - [Ollama](https://ollama.com): Run LLMs locally with a simple CLI. Wraps llama.cpp with model management, an HTTP API, and one-command model downloads. The easiest way to get started with local models. - [SGLang](https://sgl-project.github.io): Fast serving framework with RadixAttention for automatic KV cache reuse across requests. Features a frontend language for complex LLM programs with parallelism. - [TensorRT-LLM](https://nvidia.github.io/TensorRT-LLM): NVIDIA's optimized inference library for LLMs on NVIDIA GPUs. Leverages TensorRT for kernel fusion, in-flight batching, and FP8 quantization for maximum throughput. - [Text Generation Inference](https://huggingface.co/docs/text-generation-inference): Production-ready inference server for LLMs. Features continuous batching, Flash Attention, quantization support (GPTQ/AWQ), and seamless Hugging Face Hub integration. - [TokenSpeed](https://lightseek.org/tokenspeed/): Speed-of-light LLM inference engine built for agentic workloads, pairing TensorRT-LLM-level performance with vLLM-level ease of use. Features a static compiler, an optimized MLA kernel for Blackwell, and FP4 inference on NVIDIA and AMD. - [vLLM](https://docs.vllm.ai): High-throughput LLM serving engine featuring PagedAttention for efficient memory management. Supports continuous batching, tensor parallelism, and OpenAI-compatible API. ## Optimizations Overview: https://openllmstack.com/optimizations/ - [Continuous Batching](https://handbook.modular.com/inference-optimization/static-dynamic-continuous-batching/): Schedules requests at the iteration level instead of waiting for a full static batch, so finished sequences are evicted and new ones join immediately. Keeps the GPU saturated and delivers large throughput gains under concurrent load. - [FlashAttention](https://github.com/Dao-AILab/flash-attention): IO-aware exact attention algorithm that reduces memory reads/writes by tiling and recomputation. Provides 2-4x speedup and enables longer context lengths without approximation. - [KV Cache Offloading](https://handbook.modular.com/inference-optimization/kv-cache-offloading/): Moves KV cache blocks from GPU memory to CPU RAM or local storage when they are not actively needed, then loads them back on demand. Frees up scarce GPU memory to support larger batches and longer contexts. - [PagedAttention](https://handbook.modular.com/inference-optimization/pagedattention/): Virtual memory-inspired KV cache management that stores cache in non-contiguous blocks to eliminate fragmentation. Enables near-zero memory waste, larger batches, and higher throughput, and is the foundation of vLLM. - [Parallelism](https://handbook.modular.com/inference-optimization/data-tensor-pipeline-expert-hybrid-parallelism/): Distributes a model and its computation across multiple GPUs and nodes. Data, tensor, pipeline, expert, and hybrid strategies trade off memory, communication, and throughput to serve models too large for a single device. - [Prefill-Decode Disaggregation](https://handbook.modular.com/inference-optimization/prefill-decode-disaggregation/): Splits the compute-bound prefill phase and the memory-bound decode phase onto separate GPU pools so each can be scaled and optimized independently. Reduces interference between the two phases and improves both latency and throughput. - [Prompt Caching](https://handbook.modular.com/inference-optimization/prefix-caching/): Reuses the KV cache of shared prompt prefixes across requests so common system prompts and few-shot examples are computed only once. Dramatically cuts time-to-first-token for workloads with repeated context. - [Quantization](https://openllmstack.com/optimizations/quantization/): Reduces the numerical precision of weights and activations (for example, FP16 to FP8 or 4-bit) to shrink memory footprint and sometimes speed up inference. Pick a format that your hardware and serving engine support efficiently, then choose the lowest bit width that passes evaluation on your workload. On Hopper or Blackwell GPUs, a validated FP8 checkpoint is often a strong production default because tensor cores accelerate FP8. On Ampere or older hardware, 4-bit weight-only AWQ or GPTQ can make a model fit, with quality and speed that depend on the model and kernel. On a laptop or Apple Silicon, GGUF through llama.cpp or Ollama is the usual path. - [Speculative Decoding](https://arxiv.org/abs/2211.17192): Uses a smaller draft model to generate candidate tokens that the larger model verifies in parallel. Achieves 2-3x faster decoding without any quality loss. ## Agentic frameworks Overview: https://openllmstack.com/agents/ - [Agent Development Kit](https://google.github.io/adk-docs/): Open source toolkit from Google for building, evaluating, and deploying AI agents. Model-agnostic and optimized for Gemini, with multi-agent orchestration, tools, and tight Vertex AI integration. - [Agno](https://www.agno.com): High-performance, full-stack framework for building multi-agent systems with memory, knowledge, and reasoning. Model-agnostic with built-in tools, structured outputs, and a path to production (formerly Phidata). - [AutoGen](https://microsoft.github.io/autogen/): Multi-agent conversation framework enabling agents to chat with each other to solve tasks. Supports customizable agents, human participation, and diverse conversation patterns. - [CrewAI](https://www.crewai.com): Role-based multi-agent framework where you define agents with specific roles, goals, and backstories. Agents collaborate to complete complex tasks through delegation and tool use. - [LangGraph](https://langchain-ai.github.io/langgraph/): Framework for building stateful, multi-agent applications as graphs. Supports cycles, persistence, human-in-the-loop, and streaming. The agent orchestration layer of the LangChain ecosystem. - [LlamaIndex](https://www.llamaindex.ai): Data framework for building LLM and agent applications over your own data. Provides data connectors, indexing, retrieval, and agent workflows for production RAG and knowledge assistants. - [OpenAI Agents SDK](https://openai.github.io/openai-agents-python/): Lightweight Python SDK for building agentic AI apps. Features agent handoffs, guardrails, tracing, and tool integration with a minimal abstraction layer over the OpenAI API. - [Pydantic AI](https://ai.pydantic.dev): Agent framework built on Pydantic for type-safe AI applications. Structured outputs, dependency injection, and model-agnostic design from the team behind Pydantic and FastAPI. - [smolagents](https://huggingface.co/docs/smolagents): Minimalist agent library focused on code agents that write and execute Python. Simple API with tool calling, multi-step reasoning, and tight Hugging Face Hub integration. ## About - [About OpenLLMStack](https://openllmstack.com/blog/about-openllmstack/): What the site covers and how entries are sourced. - [All models](https://openllmstack.com/models/): Browse and compare every model on the site. - [Sitemap](https://openllmstack.com/sitemap-index.xml): Machine-readable index of all pages.