Principal Engineer

GenAI quality, scored in production

Meta

Meta's GenAI for Customer Support owners needed to assess feature rollouts, model releases and prompt changes across hundreds of thousands of conversations. USQ's Principal Engineer built the data models, orchestration and governance that turned third-party HumanEval results into comparable global quality scores. When a release materially regressed against its baseline, product owners could switch serving back to the previous version.

Independent conversations evaluated
100,000s
Changes assessed
Feature / model / prompt
Release comparison
Global scores
Response to a material regression
Blue-green switch

The problem

Product owners needed a reliable way to assess whether a feature rollout, model release or prompt change improved customer-support responses. Third-party evaluators reviewed hundreds of thousands of independent conversations, ranking answer quality and flagging inconsistencies.

Without a standardised view, comparing releases at that scale required manual analysis. It was also difficult to spot evaluation teams whose HumanEval output was not meeting the required standard.

Evaluation platform

USQ's Principal Engineer designed the conceptual and logical data models, then built and scheduled the pipelines in Meta's internal orchestration platform. The work included the metadata and governance needed to operate the evaluation system.

The pipeline joined response metadata and context with third-party quality annotations, then standardised and aggregated the results into global scores for release comparison.

Release decisions
Product owners could compare a release with its baseline. Where the regression was material, they could route serving back to the previous version through a blue-green switch.
Evaluation quality
The same data made it possible to identify third-party evaluation teams whose HumanEval results were below the required standard.
PII controls
PII masking and randomised ID assignment were applied across pipelines, tables and dashboards for the personal-data requirements of the work.

Python · Dimensional modelling · GenAI evaluation · PII controls · Presto

Bring one workflow. We will tell you how we would approach it.

Book a 30-minute call