Evals (Evaluations)
Systematic, repeatable tests of AI system quality: a set of real task examples with known good outputs, scored automatically or by an LLM judge, run on every change. Evals are what separate a demo from a deployable system, and building them with the customer is a core Forward Deployed Engineer practice.
Want the full picture? This term comes from GenAI for Business, a complete free book on generative AI strategy and implementation by Prof. Shubin Yu (HEC Paris).