The ICO expects a data flow mapping exercise for data flowing in, around and out of your systems; privacy vendors sell tools for a privacy team. The ones I read skip how to find the data in an app that an AI builder wired up. A data map GDPR reviewers accept comes from tracing one real user through every table, bucket, log and provider; I’d expect about ten rows.

A data map GDPR reviewers accept: what a personal data inventory is

A GDPR data map is a personal data inventory: one row per place personal data lives or goes, with 8 columns: what data, whose, why, where, who can access it, which provider receives it, how long it is kept, and its class. It feeds the Article 30 record.

The map sits underneath the other records in how to manage SaaS data compliance: agreements, deletion paths and consent all start from knowing where the data is. A data inventory is the record of what personal data you hold, where it sits and why; at this size, a data inventory and a data map are the same record in my reading, so this page uses the two names for one table. These are the columns, as my working rule:

ColumnWhat it recordsWho asks for it
What dataThe fields: email, name, phone, IP address, uploaded files, message textA user asking for a copy of their data
WhoseThe people the data is about: your users, your customers’ staff, your own teamA customer’s data processing agreement
WhyThe purpose, and the lawful basis you rely on for itYour record of processing activities
WhereThe system that holds it and the region it sits inA security questionnaire
Who can access itThe roles and named people who can read itA security questionnaire
Which provider receives itThe outside company that stores or processes it for youA customer checking your sub-processors
How long it is keptThe retention period, and what enforces itAnyone asking you to delete something
Its classPublic, internal, confidential or restrictedWhoever decides what may go into logs and prompts

The law here is the EU GDPR (Regulation (EU) 2016/679) and the UK GDPR, which the Information Commission regulates (the ICO until 30 September 2026; its guidance still sits on ico.org.uk). The regulation’s text does not use the words “data map”. The duties a data map supports are the record of processing activities in Article 30 of the EU GDPR, which must be kept “in writing, including in electronic form”, and the accountability principle in Article 5(2): the controller “shall be responsible for, and be able to demonstrate compliance with” the principles in paragraph 1. The ICO’s data mapping expectations list this as a control measure: “A data flow mapping exercise is undertaken to document the data that flows in, around, and out of information processing systems or services.” The same page carries a note that, after the Data (Use and Access) Act, the guidance is under review and may change.

Nobody asks for data mapping for GDPR in the abstract. The request arrives as a customer’s contract, which under Article 28(3) sets out “the type of personal data and categories of data subjects”, and reading one starts with what a DPA agreement is. It also arrives as a security questionnaire, a user’s access request or a breach assessment. This page is practical guidance on the engineering side, not legal advice.

RoPA in data privacy: the record of processing activities the map feeds

RoPA stands for record of processing activities, the record GDPR Article 30 requires of controllers and processors, subject to a conditional carve-out in Article 30(5). The data map is organized by where data sits and the RoPA by why it is processed, so the same rows feed both. The ICO expects the record to be informed by data flow mapping.

Article 30(1) asks each controller to “maintain a record of processing activities under its responsibility”, and Article 30(2) asks each processor for “a record of all categories of processing activities carried out on behalf of a controller”. The record’s required contents, and the exemption for organizations with fewer than 250 people with its three exceptions, are read clause by clause in the small-company carve-out in Article 30(5). I won’t repeat them here.

What a RoPA is, in practice, is the same rows sorted a different way. The ICO’s audit toolkit puts it as a control measure: “The Record of processing activities (ROPA) includes details of all processing, informed by data flow mapping exercises.” For UK organizations, the ICO’s documentation templates are “two basic templates to help you document your processing activities; one for controllers and one for processors”, and the ICO adds: “Using these templates is not mandatory.” Its advice on where to begin: “A good way to start is by doing an information audit or data-mapping exercise to clarify what personal data your organisation holds and where.”

That is the whole RoPA definition for privacy work. The meaning of ropa changes outside it: in Spanish the word means clothing, and finance uses the letters for something else.

DPIA meaning, and when a privacy impact assessment under GDPR is required

DPIA means data protection impact assessment: an assessment the controller carries out before processing that is likely to result in a high risk to people’s rights and freedoms. GDPR Article 35 names 3 cases that require one, and the regulators’ guidelines list 9 criteria: in most cases, a controller can consider processing meeting two of them to require one.

Article 35 of the EU GDPR applies where a type of processing, “in particular using new technologies”, “is likely to result in a high risk to the rights and freedoms of natural persons, the controller shall, prior to the processing, carry out an assessment”. The UK GDPR’s Article 35(1) reads the same. Article 35(3) names three cases where one is required in particular:

  • A systematic and extensive evaluation of personal aspects based on automated processing, including profiling, on which decisions with legal or similarly significant effects are based.
  • Processing on a large scale of the special categories of data in Article 9(1), or of data about criminal convictions and offenses.
  • Systematic monitoring of a publicly accessible area on a large scale.

Outside those three, the regulators’ DPIA guidelines do the sorting. The Article 29 Working Party wrote them (WP248 rev.01), and the EDPB endorsed them in its Endorsement 1/2018. They list nine criteria: evaluation or scoring; automated decision-making with legal or similar significant effect; systematic monitoring; sensitive data or data of a highly personal nature; data processed on a large scale; matching or combining datasets; data concerning vulnerable data subjects; innovative use or application of new technological or organizational solutions; and processing that prevents people from exercising a right or using a service or a contract. Then: “In most cases, a data controller can consider that a processing meeting two criteria would require a DPIA to be carried out.” They add that in some cases one criterion can be enough.

“Privacy impact assessment” is the general and older name for this kind of assessment, in my reading; under GDPR, a privacy impact assessment is the DPIA, and the FAQ below covers the difference. For a small SaaS, my reading is that an ordinary booking or CRM app may meet none of the criteria, while an app that profiles people or processes health data at scale should work through the list row by row. The data map is the first input to any data privacy impact assessment, because the description of the processing starts from its rows. Where one looks required, the ICO’s DPIA guidance includes a sample template (“You can use or adapt our sample DPIA template if you wish”), and the judgment on whether your processing is high risk is a lawyer’s to make.

What goes wrong without it

Four questions land on a founder sooner or later, and each one needs the map to answer:

The questionWho asks itWhat happens with no map
”Send me everything you hold about me.”A user making an access requestThe export covers the main database and misses the email provider and the logs
”Which data do you process, and which sub-processors touch it?”A business customer’s data processing agreementThe annex gets filled from memory, and a vendor is left off
”What was in the affected system?”You, assessing a breachNobody can say whose data or which fields were exposed
”What did the old vendor hold?”You, replacing a vendorCopies stay behind with a company you no longer pay

The second row is a contract question: under Article 28(2) a processor “shall not engage another processor without prior specific or general written authorisation of the controller”, so the customer approves the outside companies you use, and the provider column is where that list comes from. Knowing what a sub-processor is tells you which rows belong on it.

In an AI-built app, the owner may not know where customer data is stored, because the builder chose the analytics, error tracking and email services and wired them in. Each of those is a place personal data goes. At least 5 of the 21 third-party apps I audited exposed PII or PHI. I picked those 21 apps and audited them in June and July 2026; they are a selected set, not a random sample, so the count describes them and is not a rate for AI-built apps in general.

Retention runs into the same wall. You cannot set or enforce a period for data you have not found, which is why a data retention policy for a small SaaS starts from the inventory.

How to map personal data in an app: the trace, the template, the classes

The work happens in four steps, in this order: trace one user, write the rows into the template, give each row a class, then audit the rows.

Trace a user record across systems: the eight stops

Tracing a user record across systems means following one real user through 8 stops: the auth record, database tables, file storage, logs, the email provider, the payment provider, analytics and error tracking. Backups and model providers are the two places easiest to miss.

To perform data mapping on an app you did not write, pick one real user, or seed a test user, and follow them through each stop. The stack names below are examples; what any provider stores is whatever its own docs say, so check there rather than trusting a summary.

  1. 01 The auth provider's user record: open the user list in your auth service's dashboard and note every field and metadata value it holds for that user.
  2. 02 Every database table with a user id or an email column: run the schema search below and open each table it lists.
  3. 03 File storage paths: the buckets or folders where uploads land, and whether the path or the file name carries the user's id or name.
  4. 04 Application and platform logs: your own log lines and your host's request logs, searched for anything printed from a request body or a user object.
  5. 05 The email provider: its contact lists and its message logs for sign-in links, receipts and notifications.
  6. 06 The payment provider's customer object: whatever your checkout code passes in when it creates the customer.
  7. 07 Analytics and session replay: the events, user properties and identify calls your front end sends.
  8. 08 Error tracking: the user context and request data attached to each captured error.

Then the two that slip past: backups, which copy the database rows the trace found, and the model provider, which receives whatever your prompts include. If prompts carry names or emails, the next job is how to redact personal data before LLM calls.

The schema search lists columns whose names look personal. It reads information_schema.columns, whose table_schema, table_name, column_name and data_type fields name each column and its type, and ~* matches a pattern case-insensitively:

-- candidate personal-data columns; run as a role with access to every table
select table_schema, table_name, column_name, data_type
from information_schema.columns
where table_schema not in ('pg_catalog', 'information_schema')
  and column_name ~* '(user_id|email|name|phone|address|ip)'
order by table_schema, table_name, column_name;

Run it as a role with access to every table, because PostgreSQL’s docs say of that view: “Only those columns are shown that the current user has access to (by way of being the owner or having some privilege).” The pattern over-matches on purpose: name also catches filename, and ip catches description. Strike what is not personal, and keep the rest as rows.

One medical-advice app I audited wrote patient vitals and its AI provider key to its server logs, a case told in full in how to prevent insufficient logging and monitoring. The lesson I take from it: a log line is a second copy of whatever it prints, so the logs get their own row in the map.

My method for the provider stops is to list them from the package manifest and the environment variable names, not from memory: a key in the environment means data goes somewhere. Cookies and trackers set in the browser follow their own rules, the cookie consent requirements. For a Supabase project, where the data sits by region is part of whether Supabase is GDPR compliant.

The GDPR data mapping template: the columns and a filled example

A GDPR data mapping template for a small SaaS is one table: a row per system and purpose, the 8 columns above, and an owner and a date for each row. The filled example below has 10 rows. Copy it into a spreadsheet; at this size a tool is not needed.

Every cell below is an example for an imagined small SaaS on a common stack. The provider names are examples, not a statement of what any provider stores, and the lawful bases are examples, not legal conclusions.

What dataWhoseWhy (purpose; lawful basis)WhereWho can access itProvider (example)How long it is keptClass
Account record: email, name, password hash, sign-in timesUsersRun the account; contractAuth tables, EU regionBoth foundersSupabaseUntil the account is deletedRestricted (credentials)
Profile: phone, address, avatarUsersFeatures the user asked for; contractProfiles tableBoth founders, supportSupabaseUntil the account is deletedConfidential
Uploaded filesUsersThe feature they were uploaded for; contractStorage bucket, path holds the user idBoth foundersSupabaseUntil the user or the account deletes themConfidential
Application logs: user ids, IP addresses, error textUsersDebugging and security; legitimate interestsHost log streamBoth foundersYour hostPer the provider’s docsConfidential
Transactional email: address, name, message contentUsersSign-in links and receipts; contractEmail provider logsBoth foundersResendPer the provider’s docsConfidential
Payments: name, email, billing detailsPaying customersBilling; contractPayment provider customer objectOne founderStripePer the provider’s docsRestricted (payment data)
Product analytics: user id, events, deviceUsersProduct decisions; consentAnalytics projectBoth foundersPostHogPer the provider’s docsConfidential
Error tracking: user id, request dataUsersFixing errors; legitimate interestsError tracker projectBoth foundersSentryPer the provider’s docsConfidential
Support inbox: email, message text, attachmentsUsersAnswering support; contractShared inboxBoth founders, supportYour email hostSet in your retention scheduleConfidential
Backups: copies of the first two rowsUsersRecovery; legitimate interestsHost backup storageBoth foundersSupabasePer the provider’s docsSame as the rows it copies

An app that sends prompts adds one more row for the model provider, with OpenAI as an example. Paste the table into a sheet, one row per system per purpose; as a data mapping template, Excel or Google Sheets does the job, and there is no download here because the table is the template. Add two columns for the owner and the date each row was last checked.

The checklist version of the GDPR data mapping template is two rules: every stop from the trace has a row, and every row has an owner and a date. For UK GDPR work, the ICO’s documentation templates above are the official starting point, and its audit framework offers trackers as “a downloadable version of each toolkit”, so a UK data audit template does not have to start from nothing. I’d save a GDPR data mapping tool that scans systems for you for a larger estate than a two-person team’s.

Data classification scheme: the column, and the policy template that holds it

A data classification scheme for a small SaaS needs four classes, as my working rule: public, internal, confidential and restricted. Ordinary personal data is confidential. GDPR special categories, payment data, credentials and government ids are restricted. The policy fits on one page: the definitions, a handling rule per class, an owner.

The four types of data classification here are my working rule, not a GDPR requirement. The restricted class starts from Article 9’s special categories: personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person’s sex life or sexual orientation.

ClassDefinitionExamples from the mapHandling rule
PublicMeant for anyone to readMarketing pages, published docsNo limits
InternalNot personal, not for outsidersAggregate metrics, runbooksTeam only; never posted publicly
ConfidentialOrdinary personal dataNames, emails, profiles, uploaded files, usage eventsNamed roles only; sent only to providers in the map; logs carry an id, not the value
RestrictedArticle 9 special categories, payment data, credentials, government idsHealth details, password hashes, API keys, billing detailsFewest people; no new provider without review; never in logs, prompts or error payloads

The examples column ties the data classification policy to your own rows instead of generic categories. As a data classification policy template, the table plus four lines is the whole document: the four definitions, the handling rules, who classifies a new row (whoever adds it to the map), and when the classes are reviewed (with the map). There is no data classification policy PDF to download; the table prints to one from any browser. Classification is a column in the map, not a separate project. If a buyer measures you against MVSP, the restricted class is the first place to look.

How to conduct a data protection audit with the map in hand

A data protection audit at this size walks the map row by row with 6 questions: purpose and lawful basis, still needed, who has access, provider agreement, retention enforced, export and delete possible. The report is the map plus a findings column and a date.

This is my method, and as a data audit checklist it runs against every row:

  1. 01 Is a purpose and a lawful basis written down for this row?
  2. 02 Is the data still needed for that purpose, or could the row be deleted?
  3. 03 Who can access it, and should each of them?
  4. 04 Is there an agreement in place with the provider that receives it?
  5. 05 Is the retention period enforced by something that runs, not only written down?
  6. 06 Can this row's data be exported and deleted for one user?

The last question is tested with one seeded user and the data subject request handling checklist. The data protection audit report needs nothing beyond the map itself, a findings column beside each row, and the date you walked it. The ICO’s audit framework is the UK regulator’s own toolkit for this: it “will help you assess your own compliance with some of the key requirements under data protection law”, and following it “does not guarantee that your processing meets all the legal requirements that apply to you”. What I describe here is a self-check, not an audit opinion; a data protection compliance audit by an outside firm or by the regulator is a different exercise.

How to verify it

A data map is verified by a second trace: seed one user with distinctive values, use every feature, then search the database, storage, logs and each provider for those values. Every place they appear must already be a row. Anything new means the map was incomplete.

Five checks, each with the evidence worth keeping:

  1. 01 Seed a second user with distinctive values, such as a unique email address and an unusual name, and use every feature once. Evidence: the user's id and the date.
  2. 02 Search for those values everywhere: every text and JSON column in the database, not only the columns the name search found; the storage paths; the logs, while your host still keeps them; and each provider's dashboard. Evidence: a hit list.
  3. 03 Match every hit to a row. Anything new becomes a row, and the map's date moves. Evidence: the diff of the map.
  4. 04 Run the reverse: every row's system still exists and still receives data. Evidence: a note per row.
  5. 05 Compare the provider column with the environment variable names and with your sub-processor list. Evidence: the three lists side by side.

For the database search, use information_schema.columns again and filter on data_type for text and JSON types; searching only the columns the map was built from would re-find what the map already holds. Don’t assume the host kept last week’s logs: search them the same day you seed the user. My working rule is to re-run the check after any new integration, and at least once a year.

On the Production Hardening Sprint, deliverable 12.1 is verified this way: we trace a sample user through storage, logs, and providers and compare it with the inventory. The picture version of the same flows is a data flow diagram, one of the views in how to create an architecture diagram.

Where the sprint does this

On the sprint, we document what personal data is stored, where it lives, why it is collected, and who can access it (deliverable 12.1). That record goes into the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Legal advice and certification are separate services; the sprint implements and documents the technical data-handling controls. Both deliverables are listed with their verify steps in the published scope.

Common questions about mapping personal data

What is the difference between a ropa and a dpia?

A RoPA is the standing record Article 30 describes, covering the processing under your responsibility; a DPIA is the assessment Article 35 requires before a type of processing that is likely to result in a high risk to people’s rights and freedoms. Both draw on the data map: the RoPA for its rows, the DPIA for the description of the processing it assesses.

What are the differences between a PIA and a DPIA?

A DPIA is the assessment GDPR Article 35 names; “privacy impact assessment” is the broader name used outside that regulation, in my reading. The legal duty in Article 35 attaches to the DPIA, so a PIA written for another purpose counts under GDPR only if it does what Article 35 asks.

Yes, under the EU and UK GDPR, when a type of processing is likely to result in a high risk to the rights and freedoms of natural persons, and it must be done before the processing starts. The three cases listed in Article 35(3) require one in particular; for anything else, the test is whether the processing is likely to be high risk, which the regulators’ nine criteria help decide.

Who fills out a DPIA?

The controller carries it out, and must seek the advice of its data protection officer where one is designated. The ICO adds that you can decide who in your organization does the work, and that “You can outsource your DPIA, but you remain responsible for it.”

What tool is used for data mapping?

A spreadsheet holding the template on this page is enough for a small SaaS, in my reading. Discovery tools that scan systems automatically suit larger estates with many databases and teams; in an app with a handful of providers, the trace finds the same places.