Guides

How to Test Prompt Quality

Create a small evaluation set and score prompts for accuracy, completeness, format, and practical usefulness.

Define the task Write the intended user, input range, prohibited behavior, and success criteria before editing the prompt. Without a target, prompt improvement becomes taste.

Build an evaluation set Include ordinary cases, edge cases, incomplete inputs, and known failures. Keep it stable so versions can be compared.

Use multiple measures Score correctness, source support, completeness, instruction following, format validity, and actionability. Some checks can be automated; subjective criteria need a clear rubric.

Record the environment Save prompt version, model, model settings, tool access, test date, and representative failures. Retest after important provider changes.