A small, reproducible harness that runs the same model on the same task three ways (full-context stuffing, a lexical prefilter, and governed metadata) and reports real token counts and F1 for each, on your own endpoint.