Resources — Case Studies

Creating Market-Native Scam Detection Data for Real-Time Mobile AI Protection

Teaching AI when to warn users and when to stay out of the way, using market-native mobile messages across 14 languages.

Detail of teenagers watching with their mobile phones

0

Scam

0

Normal

0

Rejected submissions

Challenge

For enterprise AI teams building real-time scam protection, it’s not just about simply identifying “bad” messages. It is teaching models to recognize fraud as it actually appears on people’s phones, across languages, markets, platforms, formatting styles, and local trust signals.

A global AI platform company needed real-world mobile message data to support scam detection, model training, classifier evaluation, fraud taxonomy development, safety benchmarking, and international expansion.

The dataset needed to include both fraudulent and legitimate messages because the system had to learn not only when to warn users, but when not to interfere with normal communication.

That distinction is harder than it looks.

A scam may impersonate a bank, delivery service, government agency, employer, charity, investment platform, or other trusted institution. But legitimate messages can look remarkably similar, especially when they include payment reminders, delivery updates, insurance notices, customer support exchanges, healthcare alerts, travel confirmations, sales offers, account notices, or government awareness messages.

The work required market-native judgment, not generic translation. Scam tactics, trusted institutions, dialects, link conventions, sender patterns, and platform behavior differ by country and region. A message that appears suspicious in one market may be routine in another.

The collection also had strict data-quality requirements. Submissions had to be standalone, incoming, real-world mobile messages preserved as they appeared on a phone. They could not be AI-generated, synthetic, duplicated, near-duplicated, personal conversations, expected password-reset codes, or messages outside the defined scope.

Why This Mattered to the Technical Program Team

For the client’s technical program team, the challenge was not only sourcing messages. It was turning an ambiguous AI safety objective into a program that could be executed consistently across markets.

That meant defining what qualified, coordinating contributors across languages and regions, applying consistent classification and exclusion rules, controlling duplicates, preserving message fidelity, and delivering structured data that downstream AI and product teams could use.

PPH took ownership of that operational complexity and turned it into a repeatable multilingual data program.

PPH Approach

Productive Playhouse designed a structured collection and quality-control workflow around four core components.

Market-native collection design

Contributors provided real-world messages from their own countries and regions, helping ensure the data reflected local dialects, cultural nuance, fraud tactics, link formats, trusted institutions, platform behavior, and everyday communication patterns.

The result was data native to the markets where the system would operate, rather than English-language examples translated into other languages.

The approach also recognized that language and market are not interchangeable. English-language data was collected separately for Australia, Canada, Singapore, the United Kingdom, and India. Spanish-language coverage distinguished between Mexico, Spain, Central America and Latin America.

Scam-versus-normal classification framework

PPH created contributor guidance for distinguishing genuine scams from legitimate messages and ordinary spam.

Scam messages were identified through signals such as deception, impersonation, manufactured urgency, malicious or suspicious links, requests for personal information, and attempts to obtain money or account access.

Normal messages included legitimate communications from businesses, government agencies, healthcare providers, charities, customer support teams, travel companies, insurers, retailers, delivery services, survey providers, and other service organizations.

This distinction was central to the project. A useful scam-detection system must recognize harmful messages without repeatedly flagging legitimate communications that use similar wording, formatting, or calls to action.

Quality controls for real-world signal

The collection instructions required original, non-synthetic examples and excluded duplicates, near-duplicates, personal messages, expected one-time codes, and other submissions that would weaken or confuse the dataset.

Contributors were also encouraged to check archived and deleted message folders, helping surface authentic scam and spam examples that might otherwise have been lost.

The program recorded 2,145 rejected or duplicate submissions, demonstrating the quality-control effort required to produce the final accepted dataset.

Mobile-format fidelity

PPH required contributors to preserve messages as they appeared on their phones, including wording, formatting, links, sender patterns, and available platform context.

The collection framework supported sources such as SMS/RCS, WhatsApp, Google Messages, Google Chat, LINE, Facebook Messenger, Instagram Chat, Twitter/X Chat, Signal, Telegram, and Reddit Chat.

This fidelity mattered because scam detection can depend on small contextual signals, including URL structure, spoofed authority, urgency language, unusual formatting, sender presentation, and platform-specific conventions.

Program Results

The dataset was balanced almost evenly between scam and legitimate messages, helping the client evaluate both scam detection and the risk of falsely flagging normal communication.

The near-even overall split gave the client a strong foundation for evaluating both sides of the classification problem: identifying fraudulent messages while reducing unnecessary warnings on legitimate communications.

Coverage included Australia-English, Canada-English, Singapore-English, U.K.-English, Germany-German, Brazil-Portuguese, Mexico-Spanish, India-English, Japan-Japanese, South Korea-Korean, France-French, Indonesia-Bahasa Indonesia, Spain-Spanish, and Latin American Spanish.

The volume and scam-to-normal mix varied by locale. The dataset was balanced overall rather than artificially forcing an identical distribution within every individual market.

Final Output

The final package included an annotated dataset and post-model results. PPH also provided market-level executive summaries covering observed patterns and trends, recurring message themes, cluster analysis, and recommendations for product improvement.

Impact

The project gave the client a structured foundation for collecting and evaluating market-native scam and normal-message data across multiple languages, markets, and mobile channels.

Instead of relying on generic fraud examples, translated datasets, or synthetic prompts, PPH designed a workflow that captured how scams and legitimate communications actually appear in local user environments.

The value was not only the message content. It was the operational structure around it:

  • Market-specific collection criteria
  • Clear contributor instructions
  • Scam-versus-normal classification guidance
  • Exclusion rules
  • Duplicate and near-duplicate controls
  • Mobile-format preservation
  • Platform eligibility requirements
  • Market-level reporting and analysis

For the technical program team, that structure turned a messy real-world safety problem into a manageable, repeatable data operation.

PPH helped the client build a market-native data foundation designed to support multilingual scam detection, classifier evaluation, safety benchmarking, and international expansion.

The same framework was extended to additional markets, languages, messaging platforms, scam types, and safety benchmarks, keeping pace with ever-evolving scam tactics.

Building a multilingual AI safety or evaluation program?
Talk with our team about turning it into a structured, market-ready data operation.

Contact us

Expert linguists validate, refine, and evaluate data at every stage—ensuring AI systems perform.

Contact Us
Earth
relic
relic
relic
relic