Engine reference

How the engine works

Everything on this page is rendered from the engine's own data (RULES, DIMENSIONS,SEVERITY_PENALTY, STAGE_LABELS), so it cannot drift from the code.

Pipeline

  1. Prompt parsing parse
  2. Intent detection intent
  3. Structure analysis structure
  4. Requirement extraction requirements
  5. Ambiguity detection ambiguity
  6. Missing context detection missing_context
  7. Instruction quality instruction_quality
  8. Constraint analysis constraints
  9. Output format analysis output_format
  10. Token analysis tokens
  11. Redundancy detection redundancy
  12. Security & injection analysis security
  13. Optimization strategy strategy
  14. Prompt transformation transform
  15. Validation validation
  16. Quality evaluation evaluation
  17. Token & cost comparison comparison

Stages 1–12 run for every analysis. Stages 13–17 run when you optimize. Only parsing can fail, and only for an empty prompt. Security findings are reported but never stop the pipeline.

Rule catalogue

CodeDimensionFires when
missing_objectiveCompletenessNo task keyword, imperative sentence, or question.
too_shortSpecificityFewer than 5 content words.
vague_languageClarity"something", "stuff", "etc", "and so on".
unquantified_sizeSpecificity"short", "a few", "detailed" with no number.
unclear_referenceCompleteness"this" or "the above" with nothing included to refer to.
missing_languageSpecificityCode requested without a language or framework.
hedged_instructionClarity"try to", "if possible", "maybe".
long_sentenceClarityA prose sentence over 40 words.
negative_only_constraintsClarity3+ prohibitions and no positive requirement.
contradictory_instructionsConsistencyConflicting formats, lengths, word limits, tools, or languages.
near_duplicate_instructionConsistencyTwo sentences with ≥ 80% word overlap.
missing_output_formatOutput specificationNo output format named.
missing_length_guidanceOutput specificationContent or summary task without a length limit.
wall_of_textStructure120+ words with no sections, lists, or breaks.
duplicate_instructionEfficiencyA sentence repeated verbatim.
filler_phrasesEfficiencyPoliteness and intensifiers that do not change the task.
prompt_injection_signatureSecurityWeighted injection/jailbreak signature (critical at risk ≥ 60).
sensitive_dataSecurityCredentials (critical) or personal data (warning).
undelimited_variableSecurity{{var}} or ${var} outside tags, fences, or triple quotes.
destructive_operationSecurityrm -rf /, DROP TABLE, delete all, curl | sh.

Injection signatures version 2026.09.1. Signature matching reports risk, not proof: it can miss novel attacks and can flag text that quotes an attack to discuss it.

Scoring

  • Penalty per finding: critical 30, warning 10, info 3.
  • Overall = 100 − the sum of all penalties, clamped to 0–100, and capped at 50 if any finding is critical.
  • Each dimension = 100 − the penalties of the findings in that dimension. Structure applies only to prompts of 120+ words.
  • Grades: 85+ strong, 70–84 good, 50–69 needs work, under 50 weak.

Every report lists the checks that ran for each dimension and whether they passed, along with the evidence.

Transforms and modes

ModeTransforms
conservativeNormalize whitespace; remove exact duplicate sentences.
balanced+ Remove filler and politeness phrases (please, basically, leading I would like you to, greetings); in order to → to.
structured+ Arrange sentences into Task, Context, Requirements, Output format, and Input sections; wrap template variables in tags.

Code blocks, inline code, quoted text, URLs, XML-tagged blocks, and template variables are protected: they are masked before any transform and restored byte for byte. No transform adds content except section headings and variable delimiters, and both are reported as changes.

Candidate verification

Each mode produces a candidate. A candidate is eligible only if:

  • every protected span is still present;
  • every original sentence is still present (after the same normalization the transforms apply);
  • every extracted requirement is still present;
  • it introduces no new warning or critical finding;
  • for conservative and balanced modes, it has no more tokens than the original.

The engine then selects by goal: tokens (fewest), quality (highest score), or balanced(highest score within 115% of the original's tokens). The original is always a candidate and wins when nothing is better.

Tokens and cost

OpenAI models are counted with o200k_base (or cl100k_base for GPT-4/3.5) and labeled exact. Other providers use o200k_base as a proxy and are labeled approximate. Without a loaded tokenizer, counts fall back to characters ÷ 4, labeled heuristic. Costs use a reference price table with the date each price was checked. See the model catalogue.