MLPerf has been the standard-bearer for measuring AI system performance since its launch in 2018. During that time, MLPerf has tracked over a 100X improvement in inference performance per watt for large language models and over a 50X improvement in training speed [1] [2]. Over the past eight years, the AI industry has matured, with AI services now used daily by enterprises and consumers worldwide. Today, we are announcing an evolution of MLPerf that will make it even more useful for this maturing industry: we are launching the first version of MLPerf Endpoints

MLPerf was originally designed to help a relatively small number of cloud providers and downstream system builders inform their AI hardware purchasing decisions. Today, procuring inference compute is a salient business decision for companies of all sizes – and it means evaluating options across neoclouds, cloud providers, and managed services simultaneously. These buyers need reliable, comparable, and independent performance benchmarks to inform their purchasing decisions, and these benchmarks must be dynamic enough to keep pace with an industry that launches new models weekly. 

MLPerf Endpoints is designed to meet this more diverse and complex demand for benchmarks to inform AI system procurement decisions. Four key principles are guiding our development of enterprise-buyer-centric benchmarks:  

  1. Current: Results must keep pace with the market. Buyers can’t wait months for a benchmark round to include new hardware or models. 
  2. Comprehensive: Buyers need benchmarks that cover the many competing inference providers, systems, and workloads available for purchase. 
  3. Comparable: Buyers need to see apples-to-apples results across vendors, normalized for cost or power, to inform procurement decisions.
  4. Commentary: To support buyers in these decisions, Endpoints provides additional context for benchmark results through visualizations, data filtering, and analysis. 

MLPerf Endpoints v0.7 is a foundation release with initial results from Coreweave, Google, Intel, KRAI, and Nvidia. We want to congratulate these members on their excellent results, spanning several orders of magnitude in performance across 3 benchmarks. This is the infrastructure on which we are building a more dynamic, comprehensive, and comparable inference benchmark suite for the data center. Endpoints currently supports automated submission pipelines, continuous review tooling, and dynamic visualization of results that you can see at mlcommons.endpoints.com. We are also evolving our benchmarking rules towards more buyer-centric benchmarks.

Later this year, we will deliver MLPerf Endpoints v1.0 with more buyer-centric rules, normalization, and an expanded set of benchmarks including agentic workloads and then open the rolling submission process to our broader membership. The rolling submission process ensures Endpoints will provide current and up-to-date results that move at the pace of the market. 

We’re grateful to our 30+ supporters who have helped guide the development of MLPerf Endpoints, including AMD, Argonne National Laboratory, Broadcom, Core 42, Dell, HPE, Lambda, Oracle,  and Red Hat. Their partnership has been invaluable in stress-testing our rules, processes, and improving our infrastructure. With this release, we are just getting started. 

If you’re a system or service provider, now is the time to get involved, help shape the rules, and plan your submission at a time of your choosing. Complete this form to let us know you’re interested in submitting, and what hardware you want to submit on. The community-developed rules are actively being refined for the 1.0 release later this year, and early participants in this effort will shape Endpoints’ direction. If you’re an enterprise buyer, Endpoints is being built for you, and we want to know how we can best support you and your organization. Let us know by emailing [email protected].

We deeply value the trust our members place in us as the benchmark of record for inference. We are excited to build this future of MLPerf with all of you. 

[1] A. Tschand, A. T. R. Rajan, S. Idgunji, et al., “MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from µWatts to MWatts for Sustainable AI,” in 2025 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2025. arXiv:2410.12032.

[2] D. Kanter, M. Ahmad, H. Kassa, and S. Rishab, “MLPerf Training v4.1 Results — Press Briefing,” MLCommons, Nov. 13, 2024. [Online]. Available: https://docs.google.com/presentation/d/1KSIJBvIV9OcswF1mVbhGRN0nUWbSgbYSoxu6dHCwajM/