Resources — Case Studies

Case Study: Building a Structured Human Evaluation System for AI-Generated Video

469 videos evaluated with complete ratings · Three-level quality control · Downstream production unblocked

Stockbroker analyzing financial stock market with reflection on eyeglasses

The Client

A technology company was evaluating an AI-generated video experience that combined visual scenes, voice-over narration, and text overlays to help users explore and understand a wide range of topics.

They needed human evaluation that could answer a harder question than whether the videos technically worked: Could users trust them?

The Challenge

An AI-generated video can load correctly, look polished, and still fail the user.

It can misunderstand the prompt, teach something false, omit essential information, show visuals that contradict the narration, or deliver an experience that is technically functional but not useful.

The client needed AI-generated videos evaluated across these failure modes. The content spanned fitness, finance, history, science, technology, home improvement, health, and safety. Each video was associated with multiple versions of a user prompt, requiring reviewers to evaluate the underlying intent rather than match the content against one fixed phrase.

Reviewers needed to assess:

  •  Functionality
  • Prompt understanding
  • Factual accuracy
  • Completeness and relevance
  • Aesthetics and pacing
  • Overall video quality
  • User value

Every judgment required a written rationale. When reviewers encountered a factual claim or visual element they could not confidently validate, they were required to research it using authoritative external sources.

The real challenge was not volume alone. It was turning subjective human judgment into a consistent quality signal that a technical team could use.

The operating conditions made that harder. The client requested as many evaluations as possible by the end of the following day, but the final videos and prompts had not yet arrived. PPH needed to prepare the program before the inputs were available and mobilize reviewers quickly once the assets were delivered.

Without a controlled evaluation model, the client risked receiving opinions instead of usable records.

The Approach

1. Converted broad quality criteria into executable review logic

Rather than collapsing video quality into one subjective score, the workflow separated the major components of the experience. Reviewers evaluated whether the video worked, understood the prompt, communicated accurate information, covered the topic sufficiently, delivered a coherent multimodal experience, and provided meaningful value.

Each score was paired with a written rationale, allowing the output to show not only that a video failed, but where it failed and why.

2. Stopped unreliable evaluations before they entered the dataset

The workflow introduced reviewer-level gates at the points where continuing would produce weak or misleading data.

If the video did not load or the audio was missing, the evaluation stopped. If the prompt could not be clearly understood, the evaluation stopped. If the reviewer had no knowledge of the subject, the evaluation stopped.These controls reduced risk.

3. Created a shared definition of major and minor failure

One of the biggest risks in human evaluation is that every reviewer applies a different quality threshold.

The rubric addressed this by asking reviewers to consider two questions:

  • Does the issue prevent the user from accomplishing the goal?
  • Does the issue destroy the user’s trust in the video?

A minor visual artifact might be forgivable. Missing narration, incorrect instruction, false labeling, or visuals that contradicted the voice-over were not.

This gave the 39-person reviewer pool a common standard for distinguishing imperfections from failures that materially damaged the user experience.

4. Evaluated the complete multimodal experience

Reviewers assessed the interaction between:

  • The original user prompt
  • Spoken narration
  • Generated scenes
  • Text overlays
  • Audio and visual synchronization
  • Transitions and pacing
  • Factual claims
  • Overall usefulness

This mattered because a video could be correct in one modality and fail in another. Accurate narration paired with misleading visuals was still a major issue. A polished video that failed to answer the prompt was still not useful.

The evaluation structure made these failure modes visible rather than hiding them inside a single overall score.

5. Added factual verification to the review process

The video set included subjects where confident-sounding errors could materially mislead users, including first aid, financial guidance, exercise injuries, taxes, electrical systems, history, and science.

Reviewers were instructed to verify uncertain claims and visual representations rather than rely on assumption.

This moved the work beyond preference rating. The resulting rationales documented specific factual, visual, and instructional problems that could be acted on downstream.

6. Applied quality control at three levels

PPH uses a three-level quality-control model.

Reviewer-level gates halted evaluations where reliable judgment was not possible.

Record-level validation confirmed that every submitted evaluation was complete across all required dimensions and that each rationale met a defined evidence standard.

Delivery-level review checked consistency in how severity standards were applied across the full reviewer pool before submission.

Source-asset defects were escalated rather than absorbed into the ratings. This prevented problems with the original files from being misclassified as evaluation findings.

7. Mobilized the reviewer pool

PPH deployed reviewers to complete the final evaluation set.

The operating model allowed the team to apply a common review standard across hundreds of videos and a wide range of subject areas while maintaining complete, evidence-based records for every item delivered.

Outcome

PPH delivered complete evaluations across 469 AI-generated videos, with one structured rating record for every video in the final scope.

Each record included assessment across the required quality dimensions and a written rationale explaining the reviewer’s judgment. Record-level validation confirmed that all submitted evaluations were complete and met the required evidence standard.

The three-level quality-control process supported consistency across the reviewer pool, separated source-asset defects from actual video-quality findings, and gave the client a cleaner, more defensible human quality signal.

Most importantly, the completed evaluation set unblocked a downstream technical workstream that had been gated on human quality data.

PPH turned a broad request to evaluate AI-generated videos into a controlled program that produced specific, diagnosable feedback about factuality, prompt adherence, multimodal coherence, usefulness, and product value.

For a technical program team, PPH can operationalize complex human judgment, apply it consistently across a distributed reviewer pool, and deliver the structured quality signal required to move technical work forward.

Before and After

Before engagement

  • The client had 469 AI-generated videos but no complete human quality signal.
  • Video quality depended on factual, visual, narrative, technical, and user-value criteria.
  • Reviewers could interpret major and minor failures differently.
  • The range of subject matter increased the risk of uninformed judgments.
  • A downstream technical workstream could not proceed without completed human evaluation.

After engagement

  • All videos had a complete structured evaluation.
  • Reviewers operated against a common set of quality and severity standards.
  • Every record contained the required ratings and evidence-based rationale.
  • Three levels of quality control protected completeness and consistency.
  • Source-asset defects were separated and escalated rather than distorting the findings.
  • The completed evaluation set unblocked the downstream technical workstream.

Contact us

Expert linguists validate, refine, and evaluate data at every stage—ensuring AI systems perform.

Contact Us
Earth
relic
relic
relic
relic