Skip to content

Classified personal data.

Classify personal data by category, including special category data, with a confidence score and a source span for every finding.

The problem

Personal data is more than names and dates of birth. A sentence about someone's health, an opinion in a manager's email, a salary figure in a table. Keyword search finds what you already thought to look for and misses the rest.

How Redactics handles it

Redactics uses layered detection. Pattern matching and Azure AI Language handle the deterministic volume: identifiers, contact details, dates, references and account numbers. Contextual language models handle what needs reading: health, opinions, relationships and other special category data. Every finding is classified, scored and linked to the exact span in its source.

  • Categories aligned to UK GDPR, including special category and criminal offence data
  • Confidence score on every finding
  • Deterministic rules where rules work; language models where context is required
  • Custom categories and organisation-specific identifiers

SAR-2026-0417/Findings

Findings by category

1,428 total
  • Contact details486
  • Employment371
  • Financial118
  • Health (special category)64
  • Third-party personal data294
  • Opinions and assessments95

How findings were made

Pattern rules for identifiers and references
402
Azure AI Language for entities and contact details
611
Contextual model for health, opinions and relationships
415

1,296 at or above 85% confidence, 84 between 70 and 85%, 48 below 70% and queued for judgement.

PII detection in the guided sample case. All data is synthetic.

For your team

Start from a classified list
Reviewers begin with findings grouped by person and category, ordered by what needs judgement, rather than with a search bar.

For your risk

Special category data surfaced early
Health, ethnicity, sexual orientation and similar data is identified on the way in, not discovered on the way out.

For your security review

Models run in your region and never train on your data
Language model calls stay within Azure OpenAI in your chosen region. Prompts and outputs are not used to train any model.
Security and trust

Make SARs manageable.

Try Redactics yourself or talk to us about your current process.