Data residency requirements are rules, set by law, by a regulator, or by a contract, that control where specified data may be stored and processed. Most privacy laws do not require data to stay in one country. They restrict transfers out of it, and allow them under conditions. A smaller set of rules, such as Russia’s personal data law, India’s payment-data rule, and US tax and defense contracts, do require in-country storage for specific data.
Meeting a requirement takes three separate answers: what the rule says must stay where, how your systems enforce it, and what evidence proves they did. This guide covers each in turn.
Data residency requirements at a glance
| Jurisdiction | Must data stay in-country? | What the rule actually requires |
|---|---|---|
| EU and EEA (GDPR) | No general storage rule | Transfers outside the EEA need an adequacy decision, safeguards such as standard contractual clauses, or a narrow derogation |
| China (PIPL) | Yes, for critical infrastructure operators and high-volume processors | Exports need a security assessment, a standard contract, or certification, depending on volume |
| Russia (Law 152-FZ) | Yes, for citizens’ personal data | Since 1 July 2025, listed processing may not use databases outside Russia |
| India | Payment data: yes. Personal data generally: no | Payment system data must be stored only in India; other transfers are allowed unless the government restricts a destination |
| Saudi Arabia (PDPL) | No blanket rule | Transfers allowed for listed purposes with adequate protection or approved safeguards |
| Vietnam (PDPL) | No blanket rule | Cross-border transfers need an impact assessment dossier filed with the regulator |
| Australia | Yes, for My Health Record system records | Those records may not be held or processed outside Australia |
| United States | No general law | Sector and contract rules apply, for example to federal tax information and DoD cloud data |
The rest of this guide explains each row and the sources behind it. Laws change, so treat this as orientation and confirm the current text with counsel before committing to a design.
Residency, sovereignty, and localization
The three terms are related, and mixing them up leads to either over-building or under-complying.
- Data residency is where data is stored and processed. It is usually a requirement you meet by choosing locations and controlling transfers.
- Data localization is the strict form: specified data must stay in the country, sometimes with no copy allowed abroad.
- Data sovereignty is whose law reaches the data. Data can sit in Frankfurt and still be subject to another country’s disclosure laws through the company that operates it. Our data sovereignty guide covers that question in depth.
Residency requirements are the ones you can draw on an architecture diagram, which is why they fall to the teams that build and run data pipelines.
Legal requirements: what must stay where
European Union: transfers are regulated, storage location is not
The GDPR does not say where personal data must be stored. Chapter V instead makes any transfer to a country outside the EEA conditional. The transfer needs an adequacy decision for the destination, appropriate safeguards such as standard contractual clauses or binding corporate rules, or one of the narrow derogations in Article 49.
The transfer rules sit in the GDPR’s top fine tier: up to EUR 20 million or 4% of worldwide annual turnover, whichever is higher, under Article 83(5). That tier has been used. In May 2023 Ireland’s Data Protection Commission fined Meta EUR 1.2 billion for continuing to transfer personal data from the EU to the US.
Transfers to the US have a volatile history. The Court of Justice invalidated the EU-US Privacy Shield in July 2020 (the Schrems II judgment). The Commission adopted a new adequacy decision for the EU-US Data Privacy Framework on 10 July 2023, and the EU General Court dismissed a challenge to it in September 2025. An appeal to the Court of Justice is pending. Many organizations keep EU personal data in the EU by default, because it is simpler than defending each transfer.
China: localization for large processors, gated exports for everyone else
China’s Personal Information Protection Law requires critical information infrastructure operators, and processors handling personal information above thresholds set by the Cyberspace Administration of China (CAC), to store personal information collected in China domestically (Article 40).
The CAC’s Provisions on Promoting and Regulating Cross-Border Data Flows, in force since 22 March 2024, set the export tiers for processors that are not critical infrastructure operators, counted cumulatively from 1 January of the current year:
- Security assessment: personal information of 1 million or more people, sensitive personal information of 10,000 or more people, or any important data.
- Standard contract or certification: 100,000 or more but fewer than 1 million people, or sensitive personal information of fewer than 10,000 people.
- Exempt: fewer than 100,000 people, excluding sensitive data, plus specific cases such as transfers needed to perform a contract with the individual.
Critical infrastructure operators need a security assessment for any export of personal information or important data. Multinationals often run a separate stack inside mainland China for this reason.
Russia: an outright ban on foreign databases
Russia has required localization of citizens’ personal data since 1 September 2015. An amendment in force since 1 July 2025 made the rule stricter: Article 18, part 5 of Law 152-FZ now prohibits recording, systematizing, accumulating, storing, updating, or retrieving Russian citizens’ personal data using databases located outside Russia, with narrow exceptions. LinkedIn was blocked in Russia in 2016 for failing to comply with the original rule.
India: a sector rule today, a general framework phasing in
India’s clearest residency rule is sectoral. The Reserve Bank of India’s April 2018 circular requires payment system operators to store “the entire data relating to payment systems” only in India.
The Digital Personal Data Protection Act, 2023 takes a negative-list approach: under Section 16, the government may restrict transfers to countries it notifies, and stricter sector laws such as the RBI rule continue to apply. The DPDP Rules, notified in November 2025, phase in over 18 months. They allow transfers subject to requirements the government may set, and they let the government require Significant Data Fiduciaries to keep specified categories of personal data in India.
Middle East: Saudi Arabia
Saudi Arabia’s Personal Data Protection Law came into force on 14 September 2023, with enforcement after a one-year transition. Its transfer regulation does not require in-Kingdom storage. It allows transfers for listed purposes where the destination offers an appropriate level of protection, and otherwise requires a safeguard: standard contractual clauses, binding common rules, or a certificate of accreditation.
Asia-Pacific: Vietnam and Australia
Vietnam’s Personal Data Protection Law took effect on 1 January 2026. It allows cross-border transfers, but the transferring party must prepare an impact assessment dossier and file it with the personal data protection authority within 60 days of the first transfer.
Australia has no general localization rule, but section 77 of the My Health Records Act 2012 prohibits operators in the My Health Record system from holding those records, or processing information relating to them, outside Australia. Breach is a criminal offence.
United States: no general law, many specific ones
The US has no general law requiring commercial personal data to stay in the country. Residency obligations come from sectors and contracts:
- Federal tax information. IRS Publication 1075 states that federal tax information “must not be received, processed, stored, accessed, or transmitted to (IT) systems located offshore.”
- Defense cloud contracts. The DFARS clause 252.239-7010 requires contractors to keep government data that is not on DoD premises within the United States or its outlying areas, unless the contracting officer authorizes otherwise.
- Export-controlled technical data. Under ITAR, 22 CFR 120.54(a)(5) treats unclassified technical data sent or stored abroad as not exported only if it is end-to-end encrypted to the required standard and not stored in a proscribed country.
- Health data. HIPAA does not restrict location. HHS confirms that covered entities may use a cloud provider that stores ePHI outside the US under a business associate agreement, while noting that offshore storage may increase risk.
The Justice Department’s rule on bulk sensitive data restricts certain transfers to designated countries of concern. It limits where data may go, but it does not require data to stay in the US.
The number of laws keeps growing
According to UNCTAD’s tracker, 155 of 194 countries have data protection and privacy legislation. Not all of it touches residency, but each new law is a potential new rule for where your data may go, so design residency controls that can take on a new jurisdiction without re-architecting.
How to enforce it: architectural requirements
The legal answer tells you which data has to stay where. The architecture has to make that true for every copy of the data, including the ones nobody drew on the diagram.
1. Classify the data, not the system
Residency attaches to categories of data, so start with a short matrix before touching infrastructure:
| Data category | Example fields | What usually decides where it may go |
|---|---|---|
| Customer personal data | Name, email, IP address, device ID | Privacy law of the data subject’s location, plus contracts |
| Payment and transaction data | Card, account, transaction records | Financial regulator and payment-system rules |
| Health data | Patient records, diagnoses | Health privacy law, national health-record laws |
| Government and defense data | Contract data, tax information, technical data | Agency contract terms and export control |
| Telemetry and logs | Application logs, traces, access logs | Whatever personal or regulated data they happen to contain |
The last row is where most surprises come from. Access logs carry IP addresses, traces carry user IDs, and error messages carry whatever a developer pasted into them.
2. Map every hop, including the ones you don’t operate
For each category, write down where the data is generated, every place it transits, where it is processed, where it is stored, where it is backed up, and who can read it from where. Include vendor ingest endpoints, managed services, support access, and your own tooling’s telemetry. A log pipeline that forwards records from a German office to an observability platform hosted in another country is a cross-border transfer of whatever personal data those records contain.
3. Shrink what crosses the border
Every transfer obligation is about data that leaves. The cheapest control is to send less: drop fields that the destination does not need, aggregate where a count will do, and select the fields you keep instead of listing the fields you remove. Selection fails safe when an application starts emitting a new field; a blocklist lets it through.
4. Pin processing to the right place
Choosing a cloud region controls where one service stores data. It does not control where your pipeline processes it, where backups replicate, or where a vendor’s support engineer connects from. Residency has to be enforced where the data is transformed and routed, and the placement rule has to be written down in configuration that someone can review.
5. Keep keys and access in-region
Encrypt at rest and in transit, and hold the keys in the same jurisdiction as the data they protect. That limits who can read the data, but it does not change where the data is. Under localization rules, encrypted data stored abroad is still data stored abroad. Pair keys with access controls that stop administrators outside the region from reading raw regulated data through side paths.
How to verify it: evidence an auditor will accept
A residency control you cannot demonstrate is a residency claim. For each obligation, keep evidence that answers three questions:
- What ran where? Records of which machine or region executed each processing job, not just the intended configuration.
- What left? The actual output schema for every cross-border flow, checked against the fields you approved. Test it with a record containing an unexpected personal field and confirm the field does not appear.
- What else moved? Backups, replicas, dead-letter queues, retry buffers, debug logs, and vendor telemetry. Error paths frequently contain the full original record.
Re-run the checks when the pipeline changes. The drift that breaks residency is usually a new field, a new destination, or a new region added for capacity, not a deliberate decision.
Enforcing residency at the source with Expanso
Expanso runs data pipelines on nodes you control, at the edge, on premises, or in a cloud region, so filtering, field selection, and routing happen before anything is forwarded. Each job’s outputs are set explicitly in its configuration. Managed Expanso Cloud coordinates the fleet and receives operational metadata, including metrics, health, and logs, so include that metadata in your residency review; the security and governance page describes the boundary.
Two parts of a job map onto residency requirements:
- The selector decides where a job runs. Nodes carry labels such as
region: eu, and a job’sselectortargets matching nodes. The edge deployment docs are explicit that labels are metadata, not security boundaries, so pair them with network and access controls that stop regulated data from reaching the wrong node in the first place. - The outputs decide what leaves. A job writes only the outputs it declares, and each output can carry its own processors.
A runnable example
This job keeps the full record, personal fields included, on the node that received it, and forwards a second copy that contains only four named, non-personal fields. Save it as eu-events-job.yaml:
name: eu-events
type: pipeline
selector:
match_labels:
region: eu
config:
input:
file:
paths: [./events.jsonl]
codec: lines
output:
broker:
pattern: fan_out
outputs:
- file:
path: ./in-region/raw.jsonl
codec: lines
- file:
path: ./out-of-region/summary.jsonl
codec: lines
processors:
- mapping: |
root.event_id = this.event_id
root.country = this.country
root.action = this.action
root.duration_ms = this.duration_ms
Create events.jsonl in the same directory with two synthetic events. The second carries a customer_phone field that the job never mentions:
{"event_id":"de-001","country":"DE","email":"[email protected]","ip_address":"192.0.2.10","action":"checkout","duration_ms":412}
{"event_id":"de-002","country":"DE","email":"[email protected]","ip_address":"192.0.2.11","action":"login","duration_ms":95,"customer_phone":"+49-000-0000"}
With Expanso Edge installed, start it in local mode from that directory, then deploy the job from a second terminal in the same directory:
expanso-edge run --local
expanso-cli job deploy eu-events-job.yaml \
--endpoint http://localhost:9010
in-region/raw.jsonl now holds both events unchanged. out-of-region/summary.jsonl holds only this:
{"action":"checkout","country":"DE","duration_ms":412,"event_id":"de-001"}
{"action":"login","country":"DE","duration_ms":95,"event_id":"de-002"}
The email address, the IP address, and the phone number never reach the outbound file, and the phone number was excluded without anyone naming it. That is the verification step from the previous section, run against a real pipeline: feed it a record with a field you did not plan for and read what comes out.
Local mode runs every job on the one machine you started, so it does not exercise the selector. On a fleet managed by Expanso Cloud, the scheduler places the job only on nodes whose labels match, and you replace the two file outputs with your in-region store and your central destination. The same pattern covers PII removal from logs, and there is a runnable PII example that needs no account.
Common data residency mistakes
Treating residency as a storage-only problem
Transfer rules apply to processing and access, not only to the database at rest. EU data stored in Frankfurt but processed by a cluster in another region has still been transferred.
Ignoring telemetry and logs
Compliance reviews tend to classify customer databases and skip observability data. Logs, traces, and metrics often carry personal data: IP addresses in access logs, user IDs in trace spans, email addresses in error messages. They are subject to the same rules as the primary database. See what an observability pipeline is for where that data usually flows.
Relying on cloud region selection alone
Picking an EU region controls one service’s storage. Data can still leave through backups, replication, CDN caches, support access, and third-party integrations. Every one of those paths needs its own answer.
Assuming encryption solves residency
Encryption protects confidentiality, not location. Under localization rules such as Russia’s, encrypted data stored abroad is still data stored abroad. The ITAR carve-out for end-to-end encrypted technical data is a specific exception, not a general principle.
Treating a hash as anonymization
Hashing an identifier keeps records about the same person linkable, and small input spaces such as phone numbers can be reversed by brute force. Under the GDPR, pseudonymized data that can be attributed to a person with additional information is still personal data, so it stays subject to the transfer rules.
FAQ
What are data residency requirements?
They are legal, regulatory, or contractual rules that control where specified data may be stored and processed. Some require data to stay in a country; most restrict and condition transfers out of it.
Does GDPR require data to be stored in the EU?
No. GDPR regulates transfers of personal data outside the EEA, which need an adequacy decision, appropriate safeguards such as standard contractual clauses, or a narrow derogation. Many organizations keep EU data in the EU anyway because it avoids justifying each transfer.
Which countries require data to stay in-country?
Russia prohibits using databases outside Russia for specified processing of citizens’ personal data. China requires domestic storage for critical infrastructure operators and processors above CAC thresholds. India requires payment system data to be stored only in India. Sector rules elsewhere, such as Australia’s My Health Records Act and US federal tax rules, localize specific data.
What is the difference between data residency and data sovereignty?
Residency is where data is stored and processed. Sovereignty is whose law governs access to it, which can include the laws of the country where the operating company is based. The data sovereignty guide covers the difference in detail.
Does data residency apply to logs and telemetry?
Yes, whenever they contain personal or otherwise regulated data. IP addresses, user IDs, and free-text error messages bring observability data under the same rules as the systems that produced it.
How do I prove data residency to an auditor?
Keep records of where each processing job ran, the actual fields that left each region, and every secondary path: backups, replicas, retry queues, and vendor telemetry. Test with records containing unexpected personal fields and keep the results.
Map your residency exposure
If your data is already distributed across sites and regions, enforcing residency where the data is produced is usually simpler than centralizing it and defending every transfer. To work through a specific obligation, book a free data consultation. For the broader governance picture, see the data governance use case and the edge data governance guide.
