LLM Reference
LLM Reference is the AI command center that decodes the exploding model landscape so you can ship with the perfect provider instantly.

About LLM Reference
LLM Reference is a revolutionary decision-support directory engineered specifically for builders, architects, and technology leaders navigating the explosive, hyper-accelerated frontier of large language models. In an era where the AI landscape mutates weekly with new releases, price disruptions, and shifting benchmark rankings, this platform acts as a singular, authoritative beacon. It tracks over 1,843 models from 140 providers and 247 research labs, refreshing its intelligence daily to ensure you are never operating on stale data. The core value proposition is elegantly radical: stop wasting engineering cycles hunting through scattered, unreliable sources and start shipping with unshakeable confidence. Whether you are architecting a next-generation coding assistant, orchestrating an autonomous agentic workflow, crafting a high-fidelity writing tool, or building a sophisticated research pipeline, LLM Reference provides a single, trustworthy command center. Here, you can compare models side-by-side, instantly identify the cheapest frontier output pricing, and browse curated editors' picks for specific tasks like coding, agents, writing, research, image generation, and video creation. The interface is designed for fast triage, enabling you to identify the optimal model for your job, determine the most cost-effective provider, and return to building at lightspeed. A dynamic Pulse feed highlights weekly changes including new models, price cuts, and benchmark refreshes, keeping you informed without the noise. Built by the Data Advantage project and updated daily, LLM Reference is an indispensable resource for anyone who must stay ahead of the exploding LLM ecosystem.
Features of LLM Reference
Comprehensive Model Directory
Access a living, breathing database of 1,843 language models spanning 140 providers and 247 research labs. This is not a static list but a dynamically updated catalog that captures every significant release, from frontier giants to specialized open-weight innovators. You can search by task, provider, or capability, filtering through the chaos to find exactly what your architecture demands.
Intelligent Editors' Picks
Navigate the noise with curated, expert-driven recommendations for specific use cases. The platform features editors' picks for coding, agents, writing, research, image generation, and video creation, each backed by rigorous analysis of benchmark scores and real-world performance. These picks are not opinions but data-informed decisions, updated as the landscape shifts, showing you where most teams should start.
Real-Time Pulse Feed
Stay synced with the market's heartbeat through the Pulse feed, which aggregates weekly changes including new model releases, verified provider price cuts, and benchmark refreshes. With 177 new models, 53 price cuts, and 368 benchmark refreshes tracked this week alone, this feature ensures you never miss a critical update that could impact your deployment cost or performance.
Side-by-Side Model Comparison
Instantly compare any two models across a matrix of critical dimensions including pricing, benchmark scores, context windows, and task-specific capabilities. This feature eliminates the need for manual cross-referencing, enabling rapid, data-driven decisions. You can see exactly how a frontier model stacks against a cheaper alternative for your specific workload.
Frontier Pricing Intelligence
Track the cheapest frontier output pricing in real-time, with verified data on provider costs per 1M tokens. The platform highlights the top-lab output price, currently at $0.260 per 1M tokens via Tencent Cloud TI Platform, and monitors price cuts as they happen. This intelligence is critical for cost-conscious deployments at scale.
Use Cases of LLM Reference
Selecting a Production Coding Model
When building a coding assistant or agentic workflow, you need a model with exceptional SWE-bench performance. LLM Reference enables you to instantly identify Claude Fable 5 with its 80.3% SWE-bench Pro and 96% SWE-bench Verified scores, compare it against alternatives like GPT-5.5 or Claude Opus 4.8, and verify pricing before committing to a provider. This eliminates guesswork and reduces evaluation time from days to minutes.
Optimizing Agentic Workflow Costs
For teams orchestrating complex agent loops, the cost per call is critical. Use the platform to filter models by tool-use capability and pricing, discovering that Claude Sonnet 4.6 leads with a 87.5 tau-bench score while comparing its output cost against cheaper options like DeepSeek V4 Flash. You can then route requests to the most cost-effective provider without sacrificing reliability.
Choosing a Video Generation Provider
Creative teams evaluating video generation models can use the editors' picks to see that Veo 3.1 is rated best overall with 30-second clips and native audio up to 4K. By comparing against Runway Gen-4.5 and Wan 2.7, you can make an informed decision based on quality, resolution, and pricing, all within a single interface rather than scouring multiple provider dashboards.
Benchmarking Research Capabilities
Researchers needing a model for data analysis and trading tasks can leverage the platform's leaderboards to see that Claude Fable 5 leads with a GDPval-AA ELO of 1932. By comparing against GPT-5.5 and Gemini 3 Pro, and checking the latest benchmark refreshes, you can select the model that maximizes accuracy for your specific analytical pipeline.
Frequently Asked Questions
How often is the model directory updated?
The model directory is refreshed daily, with a dedicated Pulse feed that highlights weekly changes. This includes new model releases, verified price cuts from providers, and benchmark refreshes. The platform currently tracks 1,843 language models, 140 providers, and 247 labs, ensuring you always have access to the most current data.
What criteria are used for Editors' Picks?
Editors' Picks are based on a combination of benchmark scores, real-world performance data, and cost-effectiveness analysis. Each pick is labeled with a performance rating (e.g., Excellent) and includes specific metrics such as SWE-bench scores for coding or Chatbot Arena ELO for writing. Picks are updated as new models and benchmarks are released, reflecting the current state of the field.
Can I compare models from different providers?
Yes, the Compare feature allows you to select any two models from any providers and view them side-by-side. You can compare pricing per 1M tokens, benchmark scores across multiple suites, context lengths, and task-specific capabilities. This enables you to make informed decisions about which provider offers the best value for your specific use case.
How are price cuts verified?
Price cuts are verified by cross-referencing provider announcements, API pricing pages, and official documentation. The platform tracks 53 price cuts this week alone, ensuring that the pricing data you see is accurate and actionable. This verification process is critical for cost optimization in production deployments.
Is there a way to track changes over time?
Yes, the Changelog feature provides a historical record of all updates to the platform, including new models added, price changes, and benchmark refreshes. Combined with the Pulse feed, this gives you a comprehensive view of how the LLM landscape has evolved, enabling you to make strategic decisions based on market trends.
Similar to LLM Reference
GeoRank
Planning a relocation or long-term stay abroad? Compare places on sunshine, cost, tax, visa access for your passport, then ask AI about your short
Adviserry
Automatically get personalized actions from your YT, podcast, and email subscriptions.
AI ChatGPT Powered App
AI Chat, Chat AI, AI Detector & AI Checker. Chat with AI, detect AI writing, humanize text and create content powered by Chat GPT 5.5.
AuditBadger
SOC 2 and ISO 27001 turned into a clear to-do list. AI prepares the first drafts, you approve every call, and the founders actually answer.