We define the task and success criteria, create fixed test inputs, run the same prompt across listed models, compare completeness and failure modes, revise one variable at a time, and record the model list, version, and test date. Outputs remain examples, not guarantees.
HiddenPrompts
How we test prompts
A repeatable evaluation process—not a claim that one prompt works everywhere.