NIST AITE: When Domain Experts Grade the AI
NIST's new AITE program tests AI models on blind data defined by domain experts. The question shifts from how smart the model is to how useful it is.
In July 2026, NIST's Technology Test and Evaluation Division announced a new programme called the Artificial Intelligence Technology Evaluation - AITE. Its first evaluations begin this month. The programme tests AI models against blind data in a sequestered testbed, with initial tasks covering quantum dot control, human genome variant curation, and public safety visual event recognition, all using large vision language models. That is the news. The more consequential detail is structural: AITE does not let AI developers define the test. It lets domain experts define it. This is this article's central argument - that the design of AITE shifts who holds the authority to say what good AI performance looks like, and that shift matters more than the programme's initial scope suggests.
The design that matters
AITE's participation structure has two tracks. Data providers - quantum physicists, genomicists, public safety professionals - contribute original datasets from their domain and define a meaningful task to be performed on that data. Their datasets remain inaccessible to other participants. Model providers - AI developers - submit models for evaluation against those datasets and tasks. The evaluation data is not intended to serve as training data.
Each side gets something different. Data providers receive detailed performance measurements showing how the best-performing models handle their specific data and task. Model providers see how their models perform relative to others on the same data using identical metrics.
The structural consequence is that the people who define what a correct answer looks like are not the people who built the model being tested. A quantum scientist decides what useful quantum dot analysis means. A genomicist decides what accurate variant curation looks like. The model developer learns whether their system meets that standard - not a standard they set themselves.
Why sequestered data is not a minor detail
AITE's use of blind, sequestered data addresses a problem the AI evaluation community has been documenting with increasing specificity. A 2024 survey across six popular QA benchmarks and fifteen large language models found benchmark data contamination levels ranging from 1% to 45%, with contamination growing over time as benchmark items propagate through training corpora. Most pre-2024 benchmarks with publicly available test items carry meaningful contamination concerns by 2026.
The mechanism is straightforward: when benchmark test data appears in training datasets - whether deliberately or through the ordinary crawling of public internet text - models memorise answers rather than demonstrating reasoning. Scores rise without a corresponding improvement in capability. The result is a credibility gap between what a leaderboard says a model can do and what it actually does on data it has never seen.
AITE's sequestered testbed mitigates this directly. The data providers' datasets are inaccessible to model providers. The evaluation environment is designed to prevent train/test contamination. This is a structural countermeasure, not a policy statement - the inference here is that AITE's design treats contamination as an engineering problem to be solved by architecture, not by asking model developers to be more careful about decontamination.
A methodology NIST has used before
AITE is not NIST's first use of sequestered-data, blind-testing evaluation. The agency has conducted independent evaluations of face recognition technologies since 1999, initially through the Face Recognition Vendor Test (FRVT, now FRTE/FATE). Since at least 2006, FRVT has measured performance with sequestered data - operational datasets not available in the public domain, such as mugshots and visa application photos. Vendors submit algorithms for third-party testing, and NIST measures performance against data the vendor has never seen.
The parallel to AITE is direct: blind data, third-party administration, vendor-submitted systems, standardised metrics. The difference is that FRVT evaluated a single technology type (face recognition algorithms) against a single data modality (facial images). AITE evaluates general-purpose AI models - large vision language models, initially - against domain-specific tasks that domain experts, not AI researchers, define.
This is a structural observation, not something NIST states about its own programme's purpose. But it is worth naming: applying a proven metrological methodology to a technology category that has, until now, been evaluated almost entirely by its own developers is a significant shift in who holds the measuring instrument.
What AITE does not do
AITE is not the same programme as NIST's pre-deployment safety testing under the Center for AI Standards and Innovation (CAISI). In May 2026, CAISI announced pre-deployment testing agreements with Google DeepMind, Microsoft, and xAI, expanding its frontier model evaluation to five major labs alongside existing partners OpenAI and Anthropic. CAISI evaluations cover cybersecurity, biosecurity, and chemical weapons risks - they ask what AI systems should not do. AITE asks what AI systems can do, in specific domains, against data they have never trained on.
The two programmes answer different questions. CAISI asks: does this frontier model pose unacceptable risks before deployment? AITE asks: does this model perform well on a task a domain expert defined as meaningful? An organisation deploying AI in a domain AITE covers should expect to encounter both questions - and should not assume that a model which passes one has passed the other.
AITE's current scope is narrow: three tasks, one modality (vision language models), three domains. NIST has indicated that additional themes - including video and natural language processing - will follow. Whether the programme scales depends on whether domain-expert data providers materialise and whether model developers participate when results might show their models underperform against domain-specific tasks they did not design. No rule requires either participation, and the programme is voluntary.
What changes for an organisation deploying AI
For an organisation evaluating AI models for deployment in a domain AITE covers, the programme offers something no public benchmark currently provides: a metrologically grounded answer to the question "does this model work for our use case," produced by a third party with no stake in the answer, against data the model has never seen.
That is a different question from "is this the highest-scoring model on the leaderboard," and it is a more deployment-relevant one. An organisation deploying a vision language model for genomic variant curation can now ask whether NIST's evaluation of that model, against blind genomics data, supports the deployment decision - rather than relying on general-purpose benchmark scores that may reflect contaminated training data rather than genuine capability.
The inference is the article's own, not NIST's: AITE does not position itself as a replacement for benchmarks or a certification programme. But its design - domain-expert-defined tasks, sequestered data, standardised metrics, third-party administration - provides a credibility floor that public benchmarks, by their structure, cannot match. Whether that floor becomes the standard against which AI deployment decisions are made is a question AITE's first evaluations, starting this month, will begin to answer.
Map your disclosures against AI & Technology.
Noa reads your disclosures, traces every number to its source, and flags what's missing.