Papers › Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
Paulius Rauba, Nabeel Seedat, Max Ruiz Luyten, Mihaela van der Schaar
The predominant de facto paradigm of testing ML models relies on either using only held-out data to compute aggregate evaluation metrics or by assessing the performance on different subgroups. However, such data-only testing methods operate under the restrictive assumption that the available empirical data is the sole input for testing ML models, disregarding valuable contextual information that could guide model testing. In this paper, we challenge the go-to approach of data-only testing and introduce context-aware testing (CAT) which uses context as an inductive bias to guide the search for meaningful model failures. We instantiate the first CAT system, SMART Testing, which employs large language models to hypothesize relevant and likely failures, which are evaluated on data using a self-falsification mechanism. Through empirical evaluations in diverse settings, we show that SMART automatically identifies more relevant and impactful failures than alternatives, demonstrating the potential of CAT as a testing paradigm.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2410.24005")
Code
Syntology Ran 16 of 27 code samples harvested from 2 repositories linked to this paper; 11 have no recorded run. Of those that ran: 2 ran · honoured contract; 11 ran · our draft was wrong; 3 ran · fixture could not drive it.
By repository: official repository: 27 samples from 2 repositories, 16 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
27 samples harvested; 16 ran; 2 honoured the contract we drafted; 11 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 27 of the 27 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 2 repositories linked to this paper, official or community; each sample names its own and says which. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
980b86cc82a40ff3 · report
082775765bcf5927 · report
644cdbc5b96ea94b · report
a842e24eaf74fcd2 · report
04ef2bc634603850 · report
952dd38f279afee7 · report
48299ec78578a7ca · report
ddea085ab8189264 · report
fa2e7f5436b471ff · report
c1530c3092d8eb42 · report
a17ed6b8532eb0da · report
78bb22f3c0eb1ec3 · report
861a293ada0282e4 · report
fbfa3ff05c9f3076 · report
3f9858e6a6582447 · report
dd22a0f722bbba34 · report
d6be180787be0544 · report
352bd9c2249abd79 · report
64d7572fd4d95cd6 · report
7b95bcfb0f225d7e · report
655af21881ab2ee3 · report
e041ca87d939d79f · report
5f50e8beb718878d · report
8eafe28f193a5cb3 · report
859fdcdc54b12008 · report
dee4e53b7cfcbb6d · report
b0c4fc6d6786516a · report
Tasks
Results from the paper archive 2025-07-28
No leaderboard rows for this paper in the archive.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections