Chinese AI Adoption Is Already Mainstream

Chinese AI adoption is no longer a future prospect. It is already happening, and the shift in Chinese open-weight models winning global market share means enterprises can no longer afford to ignore them. Investment and Innovation Behind Chinese AI Adoption: Capital Still Rules, But Efficiency Wins In 2025, US private AI investment reached $285.9 billion, more than 23 times China’s $12.4 billion. The United States also introduced 1,953 new AI startups in the same year, an order of magnitude more than any other nation. Four US hyperscalers (Amazon, Google, Microsoft, and Meta) have accelerated capital expenditures, too, with Google alone reporting more than $150 billion in annual capex in 2025. However, private investment figures likely understate China’s total AI spending. Government guidance funds have deployed an estimated $184 billion into Chinese AI firms between 2000 and 2023. Chinese labs achieve competitive results at dramatically lower costs, leveraging open-weight ecosystems, efficient model architectures, and massive domestic datasets. This cost differential is the primary driver of Chinese AI adoption worldwide. DeepSeek V4 Pro charges $0.87 per million output tokens. Anthropic’s Claude Fable 5 lists at $50 per million output tokens for the same amount, a more than 50-fold difference. When enterprises can achieve comparable performance for a fraction of the cost, the business case for switching becomes overwhelming. The battleground is no longer just about who spends the most. Instead, it’s about who spends most wisely. Overall, on that measure, Chinese AI has decisively won the efficiency battle. Technical Showdown: Where Each Side Leads, and Where They Are Tied To understand the 2026 AI landscape, first, it helps to disaggregate “leadership” into its component parts. Where the US Still Leads Clearly Domain Evidence Notable AI Models In 2025, the US produced 50 major AI systems, compared with China’s 30. Hardware and Infrastructure The US hosts 5,427 data centers, more than ten times any other country. Nvidia’s GPU ecosystem and US cloud providers remain the backbone of global AI training. Basic Research and Innovation The US continues to lead in frontier model development and foundational AI research. Where Chinese AI Adoption Has Overtaken the US Domain Evidence Open-Weight Adoption and Market Share Chinese open-weight models captured 41% of Hugging Face downloads this spring, surpassing US models. On OpenRouter, Chinese models now account for as much as 63.5% of global inference share, with the top six most popular models coming from Chinese firms including Tencent, Xiaomi, DeepSeek, MiniMax, and Z.ai. Anthropic’s Claude Opus 4.7 trails in seventh place. Academic Output and Patents In 2024, China accounted for 74.2% of global AI patents granted (97,206 of 131,121), compared with just 12.1% for the US. Among the top 100 most-cited AI papers globally, Chinese institutions contributed 41, while US institutions contributed 46, a near-tie. China also contributed 17.8% of global AI research papers in 2024, more than double the US share of 7.6%. Industrial Robotics China continues to install more industrial robots than the rest of the world combined, accounting for 54% of global installations in 2024. Where the Two Sides Are Effectively Tied Domain Evidence Model Performance According to the Stanford AI Index 2026, the AI model performance gap between the US and China has “effectively closed.” Leading systems from both countries have traded top positions multiple times since early 2025. By March 2026, Anthropic’s leading model held an advantage of just 2.7%, a margin thin enough to flip on the next major release from either side. Top-Tier Contenders On the LMArena Text Generation leaderboard (July 2026), Claude Fable 5 leads at 1,507 Elo, with multiple Anthropic models in the top tier. Kimi K3 and other Chinese models rank competitively in the same evaluation framework. The data shows a clear pattern: US leadership is real but domain-specific, not absolute. For mainstream enterprise adoption, performance parity combined with cost advantage makes Chinese AI the logical choice for a growing share of workloads. Beyond the Superpowers: Europe, the Global South, and the Multipolar Reality The US-China dynamic dominates headlines, but the global AI market is becoming genuinely multipolar. This fragmentation creates openings for Chinese AI adoption to become the default choice for cost-conscious enterprises worldwide. Europe: Sovereignty Through Regulation and Partnership Paris-based Mistral AI has become Europe’s undisputed AI champion, with annualised revenue reaching $400 million in early 2026, up roughly twentyfold from the prior year. Its enterprise clients include ASML, TotalEnergies, and HSBC. Microsoft’s multibillion-dollar partnership with Mistral is particularly telling. Microsoft is renting computing capacity from data centers Mistral is building and financing in Europe. In exchange, Microsoft distributes Mistral’s “sovereign” AI models to EU-based Azure customers. Notably, this model is directly transferable to the story of Chinese AI adoption. Just as Mistral provides European sovereignty, Chinese AI can provide cost-effective sovereignty for enterprises in Asia, Africa, Latin America, and beyond. The demand for alternatives to US-centric AI is global, and Chinese providers are positioned to fill that gap. The EU AI Act has become the world’s first comprehensive AI law, establishing a risk-based framework that many governments outside Europe now look to as a template. The Global South: The Largest Growth Opportunity Notably, India’s AI Impact Summit 2026 was a landmark event, the first time a major global AI conference was hosted by a developing country. The summit concluded with 88 nations and organisations signing the New Delhi Declaration on AI, structured around seven pillars including democratising AI resources and economic growth. Under India’s BRICS Presidency in 2026, ten nations convened separately to chart a Global South roadmap on responsible AI. In particular, Southeast Asian nations emphasise applied, frugal AI instead, solving real problems in healthcare, agriculture, and citizen services before chasing frontier models. The aggregate AI-driven productivity market in emerging economies is valued at an estimated $47.3 billion, growing at 31.4% annually. Given all this, Chinese AI adoption is uniquely well-suited to serve these markets. The cost structure of Chinese models aligns closely with the budget constraints of enterprises in the Global South. The open-weight approach enables local deployment and customisation, too,
Starlink Surveillance: Why the Beam Must Know You

Starlink surveillance is not a feature anyone switched on. It’s a side effect of orbital mechanics. A satellite in low Earth orbit moves at roughly seven and a half kilometres per second relative to the ground. To deliver internet to a dish on a roof, it has to steer a narrow beam onto that dish and hold it while both are moving. It then has to hand the connection to the next satellite before the first one drops below the horizon. None of that works unless the network knows where the terminal is, continuously, to a useful precision, for every active terminal. That is not a design decision anyone made and could unmake. Rather, it’s a consequence of the physics of the service. Starlink surveillance, in the most literal sense, is therefore not optional. The largest satellite constellation ever built maintains, as a condition of functioning at all, a live position register for every one of its users. The Scale Behind Starlink Surveillance Starlink has somewhere between 9,600 and 10,400 active satellites, depending on which tracker and which date, roughly 65% of everything operational in orbit. Subscriber counts run from 10 million in February 2026 to 12 million across 164 countries by June. That latter figure comes from the company’s own S-1 filing. The FCC authorised another 7,500 satellites in January, and filings exist for close to 30,000 in total. The newer part is direct to cell. More than 650 satellites now connect to ordinary smartphones, with no dish and no special hardware required. Coverage spans 22 countries, covering a claimed 400 million people. In December 2025, Airtel Africa agreed to deploy it across 14 African markets, reaching a customer base of 174 million. Separately, several million connected vehicles report telemetry to a company under the same ultimate ownership. What Starlink Surveillance Data Actually Exists The value in this data was never going to be content. Traffic is encrypted and mostly dull. Instead, the value sits in the layer underneath, which is generated automatically and cannot be turned off without breaking the product. Specifically, a terminal has a location, an activation time, and a deactivation time. It can also move, and when it does, the network follows it, because otherwise the beam misses. Aggregate those, and you have demand density. That’s a proxy for where people and activity are, in exactly the parts of the world that produce the least conventional reporting. Nothing in that paragraph is an allegation. In fact, it is simply a description of what the system necessarily computes in order to work. Seven Starlink Surveillance Scenarios Worth Naming These are possibilities, not findings. Each is technically available given the architecture. None is evidenced here as having occurred. They are worth listing because the time to think about them is before the dependency is universal, not after. Scenarios One Through Four Scenarios Five Through Seven Where the Vehicle Data Side Is Weaker Than It Sounds The vehicle fleet is often described as a rolling camera network feeding a central system. On the published evidence, though, it is not that, and the accurate version is narrower. Tesla states that Sentry Mode and Dashcam recordings are processed and stored in the vehicle or on external media, never on its servers. The Dutch data protection authority investigated Sentry Mode. Tesla changed the defaults in response, so recording now requires the vehicle to be touched and the owner to enable it, and the regulator closed the case without a fine. Three Qualifications That Survive Video is reachable by warrant, however, even if it never leaves the car, which is exactly what the towing cases demonstrate. Telemetry is a separate category from video, too. Location, charging, and driving data do go to the company under stated conditions, and cabin camera footage is shareable if the owner enables data sharing. And the protections are policy commitments and setting defaults rather than architectural impossibilities, which means they can change with a release note. So the fleet is not the eye, then. Instead, it’s a second, independent, high-frequency behavioural dataset about a different population, under the same ownership as the first. The One Starlink Surveillance Scenario Already Documented Selective availability is not hypothetical. SpaceX geofences the service, restricting it outside Ukraine’s borders and in Russian-occupied territory. The Pentagon’s 2023 terminal contract, separately, was written specifically to give the Pentagon control over where the service worked inside the country. The Crimea episode is worth telling accurately, because the widely circulated version is wrong. In 2022, Musk declined a Ukrainian request to extend coverage to Crimea during an operation against Russian ships, citing escalation risk and US sanctions exposure. He did not switch off existing coverage, since it had never been enabled there in the first place. CNN and The Washington Post both initially reported that he had, then issued corrections once the actual sequence of events became clear. Three senators on the Armed Services Committee wrote to the Defence Secretary about it anyway, which indicates how the incident was received institutionally. Overall, the structural fact is what matters here: a decision by a private individual shaped the communications available during a military operation. The Response That Tells You the Most The United States government did not regulate its way around this. Instead, it bought a different system. Starshield is the government variant: government-owned, more heavily encrypted, restricted to authorised agencies, backed by a reported $1.8 billion National Reconnaissance Office contract and Space Force awards. Reporting on the first Starshield award noted unusual terms and conditions. Industry observers read this as language preventing unilateral withdrawal of service, regardless of how the military chose to use it. Ukraine’s 3,000 terminals were subsequently moved onto it. The customer with more leverage than anyone else looked at the commercial arrangement. It decided that ownership and control were unacceptable risks, and paid to build a parallel version it controls instead. Everybody else remains on the commercial product. Why the Ownership Question Is Not a Normal Ownership Question Infrastructure risk models
Telesurgery Latency: Why 199 Milliseconds Is Safe

A number that would ruin a video game turns out to be fine inside an abdomen. In September 2025, surgical teams in Kuwait and Brazil operated on each other’s patients across 12,034 kilometres, in both directions, over a link averaging 199 milliseconds of telesurgery latency. Two hernia repairs. Both patients fine. If you have ever played anything competitive on a bad connection, that number should worry you. Two hundred milliseconds of lag makes an interactive system feel broken. You act, the world has already moved, and your input lands in a past that no longer exists. In fact, it should not worry you here, and the reason why is more interesting than the record itself. Why Telesurgery Latency Breaks a Game and Not an Abdomen Latency becomes destructive under two conditions: the environment keeps evolving without you, and something inside your lag window is exploiting it. A competitive game has both. Everyone else is acting continuously, and at least one of them is deliberately timing actions to land inside your delay. You are not just slow. Rather, you are slow against someone optimising against your slowness. An operating field has neither. Nothing in the abdomen is trying to anticipate the instrument. The loop is closed and self-paced: move, observe, adjust. Delay the whole loop consistently, video included, and the surgeon adapts the way anyone adapts to a heavy tiller. One correction, since most write-ups skip it. Tissue is not static. It moves with the heartbeat, with breathing, with peristalsis, and much of the foundational latency research used stationary targets, which is exactly what later work criticises. Tissue moves predictably and without intent. That is a much easier problem than tissue that moves adversarially, but it is not nothing. Why 199 Milliseconds of Telesurgery Latency Was Not Luck The number was engineered to land inside a boundary that surgical research mapped a decade ago. The reference study, run on a robotic simulator called the dV-Trainer, tested surgeons at latencies from zero to a full second. It measured completion time, instrument motion, and errors. Performance degrades exponentially, not linearly, as a result. The bands that came out of it are now standard. Two hundred milliseconds and below is ideal. Three hundred is suitable. Four hundred to five hundred works, but tires you out. Six hundred to seven hundred is fine only for simple, low-risk procedures. Past 800, the advice is to stop operating and start mentoring whoever is actually in the room instead. A telesurgery latency of 199 sits right at the top of the best band. In short, that is what “considered safe” means here. Two caveats worth naming. First, degradation does not start at 200: controlled studies find measurable differences between zero and 70 milliseconds, and again between 100 and 150. Second, a 2026 review asks whether the clinical claims built on this literature are over-optimistic, largely because so much of it ran on static models with operators who knew they were being watched. Telesurgery Latency Is Mostly Not About Distance Light in fibre crosses 12,000 kilometres in about 60 milliseconds one way, and real routes run longer than that because cables follow coastlines. The Kuwait-to-Brazil path went via Marseille to São Paulo. But a large share of the delay budget is not propagation at all. It is video: capturing, encoding, sending, decoding, and displaying a high-definition stereo image. In one experimental setup, a 20 millisecond network round trip sat underneath a 50 millisecond encode-and-decode penalty. Below continental distances, in other words, the codec beats the cable. That is why the interesting work on telesurgery latency right now is video pipelines, jitter buffering, and predictive rendering, rather than shorter routes. What Actually Changed to Cut Telesurgery Latency Three things changed, and none of them were surgical. The Robot Stopped Needing a Special Network Toumai, built by Shanghai MicroPort MedBot, was designed to run over 5G, ordinary broadband, dedicated fibre, or satellite, and to support one-to-many and many-to-many connections instead of a single dedicated link. Published testing puts bidirectional latency under 50 milliseconds within a country and under 150 across continents. Europe’s first telesurgery, between two Belgian hospitals in May 2025, ran at 20 milliseconds over a hospital’s ordinary network. That is the whole difference from 2001. The Lindbergh operation, New York to Strasbourg at 155 milliseconds, needed the best dedicated line money could buy, and the cost is exactly why nobody repeated it at scale for twenty years. The current generation needs connectivity a mid-sized hospital already has. Regulation Caught Up Toumai received Chinese regulatory approval for commercial telesurgery in May 2025, the first anywhere, and it has also run in the US under an FDA investigational device exemption. Orders climbed from around 100 in late 2025 to more than 300 by mid-2026, across roughly 60 countries. Satellite Closed the Remaining Gap In 2025, a Japanese team achieved the first robot-assisted lung resection over Starlink, in a preclinical animal study: the surgeon operated from Fukuoka, and the swine subject was a thousand kilometres away in Fukushima. The procedure took two hours and forty-four minutes at roughly 130 milliseconds of telesurgery latency, with brief image disturbance about every five minutes from satellite handover or weather. The team picked satellite specifically to cut communication cost. A country without a convenient subsea cable landing is no longer automatically out. What Happens If the Telesurgery Latency Link Drops This is the first question everyone asks, and it already has an answer. The standard design is the dual console. Specifically, a second, fully capable surgeon sits beside the patient, and control can be handed over or seized. In one published series, the swap took three seconds. So the safe state is not instruments freezing on link loss. Instead, it’s a qualified human in the room who can finish the operation. Which reframes the whole thing. Telesurgery today is not surgery without a local surgeon. Rather, it’s a distant specialist borrowing the hands of a local team. That’s a more modest claim than the marketing,
Install GrapheneOS on Google Pixel: The Full Guide
This guide shows you how to install GrapheneOS on Pixel hardware, plus what the hardening actually buys you once it’s running. GrapheneOS is a hardened Android distribution that runs only on Google Pixel devices. That restriction isn’t brand preference. Pixels are the devices that let you install a third-party OS and then re-lock the bootloader with your own verified boot key. This single property is what separates GrapheneOS from most custom ROMs. Installing GrapheneOS on a Pixel takes roughly twenty minutes on a supported device. Everything in the walkthrough below follows the project’s official web installer documentation. Before You Install GrapheneOS on Pixel: What You Need Device and Software Requirements Only officially supported Pixels are covered by the project. The published verified boot key hashes currently span the Pixel 6 family through the Pixel 10 family, including the 9a and 10a. Fourth- and fifth-generation Pixels are handled differently, though, and only display the first 32 bits of the boot key hash. As a result, they can’t be verified using the method described later. You need 2GB of free memory and 32GB of free storage to install GrapheneOS on Pixel from the web installer. It runs on Windows 10 or 11, macOS Sonoma through Tahoe, several mainstream Linux distributions, and ChromeOS. It even runs on an Android phone or tablet, which surprises people who assume a desktop is mandatory. Avoiding the Most Common Install Failures Carrier variants cause the most common failure, and it happens before the install even starts. Carrier SKUs ship with a non-zero carrier ID written to the persist partition at the factory. That ID activates carrier configuration in the stock OS, including disabling both carrier unlocking and bootloader unlocking. The carrier may be able to clear it remotely, but support staff often don’t know how. So, buy a carrier-agnostic device instead. Supported browsers are Chromium, Chrome, Edge, Vanadium, and Brave. Brave needs Shields disabled, because it caps reported storage to resist fingerprinting, and the installer then has too little space to work with. Firefox isn’t supported at all, since it doesn’t implement WebUSB. Several widely shared guides say Chrome or Firefox, and that’s simply wrong. Avoid Flatpak and Snap browser builds, too. Ubuntu’s Chromium Snap ships with broken WebUSB. Don’t use a private browsing window either, since that usually starves the installer of the storage it needs to extract the release. And don’t install from inside a virtual machine, since USB passthrough is unreliable there. Use the USB-C cable that came with the device where possible. Connect directly to a rear port on a desktop, or a port on a laptop. Avoid hubs and front-panel ports, since bad cables and hubs are the single most common source of install failures overall. With prerequisites out of the way, here’s how to install GrapheneOS on Pixel hardware in nine steps. Steps 1 – 5: Prepare the Device Steps 6 – 9: Connect, Flash, and Lock Verify GrapheneOS Installed Correctly on Your Pixel Once you install GrapheneOS on Pixel hardware, verification is the step people skip and shouldn’t. Disable OEM unlocking first. The final setup screen has a toggle for this, checked by default. Leave it checked. It can be changed later in developer settings if needed. From there, two verification mechanisms exist, and they’re the reason this OS is worth the trouble in the first place. Boot key hash. When running an alternate OS, the device shows a yellow notice at boot containing the SHA-256 of the verified boot public key. Compare it against the hash published for your model on the project’s install page. Sixth-generation Pixels and later show the full hash. This confirms that what you flashed is what the project published, even if the computer you flashed from was compromised. Hardware attestation. The project’s Auditor app uses the device’s secure element to attest that hardware, firmware, and OS are genuine. Because the point is to learn about the device without trusting it to be honest, results aren’t shown on the device being checked. You need a second Android device running Auditor and a QR code exchange, or the optional monitoring service for scheduled checks with email alerts. What Installing GrapheneOS on Pixel Actually Buys You This is where GrapheneOS separates from privacy ROMs generally. Once you install GrapheneOS on Pixel hardware, most of the work is defence against exploitation of unknown vulnerabilities, rather than feature-level privacy toggles. Verified Boot Survives the Install Most custom Android ROMs require leaving the bootloader permanently unlocked. That removes verified boot and leaves the device open to persistent compromise. GrapheneOS, in contrast, flashes its own verified boot key into the secure element. Every boot then verifies the full firmware and OS image chain against it, with rollback protection tied to the security patch level. If any OS partition is modified, the device simply refuses to read the modified data. You keep the hardware security model, instead of trading it away for the OS. Memory Corruption Defences The project ships its own allocator, hardened_malloc, with fully out-of-line metadata that rules out traditional allocator exploitation. It also adds deterministic detection of invalid frees, plus zero-on-free with write-after-free detection. On top of that come randomised and deterministic quarantines, delaying reuse to blunt use-after-free bugs. Guard pages surround larger allocations and slabs for small ones, alongside random canaries and hardware memory tagging for slab allocations. The kernel is hardened alongside it, too. That means 4-level page tables on arm64, raising ASLR entropy from 24 to 33 bits. It also means memory tagging in the main kernel allocators, canaries on the kernel heap, and memory zeroed on release in both the page allocator and the slab allocator. Unused memory is also zeroed at early boot to clear anything left from a previous boot, module signing is forced, and the kernel runs in lockdown confidentiality mode. The practical effect is that whole classes of memory bugs become unexploitable or unreliable, rather than merely unpatched. Dynamic Code Execution Is Heavily Restricted The Android runtime’s JIT compiler is
Revolut Data Breach: 75 Million Records for $500

A Revolut data breach claim surfaced on July 25, 2026, when a threat actor listed a database on a cybercrime forum, advertised as 75 million customer records. Researchers at Cybernews obtained and examined the published sample. Revolut, for its part, has said it sees no indications of any breach. Nothing has been confirmed either way. That’s precisely what makes this Revolut data breach claim useful as a case study. Most security work happens in exactly this state. A claim exists, the evidence is partial, and the vendor denies it. Someone still has to decide what to do before anyone knows the real answer. The technical question here isn’t whether Revolut was breached. It’s what the artefacts themselves can actually tell you. What Is Actually on the Table The seller published a sample of just over 100 records across four CSV files, with a fifth file surfacing later. The fields reported are unusually rich. Card data includes last four digits, card type, expiry dates, and status flags such as active, blocked, and frozen. Credentials were hashed with bcrypt or argon2id, alongside timestamps recording when each password was last rotated. Identity and account data is broader still. It covers emails, full names, phone numbers, country of residence, addresses, currency, registration IP address, subscription plan, KYC status, an internal risk score, last-activity timestamps, monthly spend, and lifetime top-up. Device data includes device models, operating systems, and timestamps. The fifth file, meanwhile, reportedly contains bank account numbers, user IDs, and SWIFT codes. The newest records in the sample date to around May 2025. Researchers found no link to previously documented breaches, and noted the presence of referral programme flags. Revolut’s position, given to Cybernews and then updated after publication, is direct on this point. It checked the user and card identifiers in the alleged records against its own systems. None correspond to valid or genuine Revolut identifiers. The review is ongoing. Signal One: The Price Is Wrong The listing is advertised at $500. This is the strongest single indicator in the whole affair, and it doesn’t require access to anything. Pricing on criminal markets isn’t sentimental. Fresh, validated financial records with card data attached command real money, because they convert directly into fraud. A genuine, exclusive dump of 75 million banking records from a live institution would be priced in the tens of thousands at minimum. It would more likely be sold privately, too, rather than listed openly on a forum. Five hundred dollars, by contrast, is combolist pricing. It’s what you charge for aggregated material other people already have, for stale data, or for something you can’t substantiate. Either the seller isn’t confident in the goods, or they already know the buyer pool values them at close to nothing. Signal Two: The Number Is Too Round The second red flag in this Revolut data breach claim is arithmetic. Revolut publicly reports around 75 million retail customers. The listing claims exactly 75 million records. That match should raise an eyebrow on its own. Real dumps don’t equal a company’s total customer count. Instead, they equal whatever the attacker actually reached. That might be one table, one shard, one export, one misconfigured bucket, or one compromised admin session. The resulting figure is arbitrary and untidy as a result. That’s exactly why genuine disclosures tend to produce numbers like 50,150, not clean round ones. A record count that exactly matches a headline marketing figure is a number that was chosen, not counted. It suggests the seller reached for the company’s own public statistic to size their claim. Signal Three: The Schema Doesn’t Look Like a Bank This is the most interesting part technically, and it cuts in an unexpected direction. Look at what’s in the field list. Subscription plan, KYC status, risk score, monthly spend, lifetime top-up, last-activity timestamp, device model, operating system, registration IP, referral flags. Now look at what’s thin instead: card data limited to last four digits, type, expiry, and status. That’s not the shape of a core banking ledger. Rather, it’s the shape of an analytics or growth data mart. This is the kind of denormalised table assembled to answer questions about customer segments, activation, and referral performance. Lifetime top-up and monthly spend are aggregate metrics, not transaction records. Similarly, risk score and KYC status are decision outputs, not the underlying evidence behind them. Suppose the data is genuine. Then that schema points away from a core systems compromise, and toward something adjacent instead: a business intelligence environment, a marketing or CRM platform, a data pipeline, or a third party with a feed. Aggregated warehouses are consistently softer targets than the systems they draw from. They also routinely hold a wider slice of customer attributes than any single production service does. Suppose instead that the data is fabricated or assembled. In that case, this same schema is exactly what you’d expect from someone stitching together material from older breach compilations, then dressing it up with plausible-sounding internal fields. The schema alone doesn’t settle it. It does, however, tell you which of the two stories to test first. Signal Four: What the Identifier Check Proves, and What It Doesn’t Revolut’s most substantive technical statement, in this Revolut data breach claim, is that the user and card identifiers in the sample don’t correspond to valid or genuine identifiers in its systems. That’s a real check, it’s fast to run, and it’s the right first move. When the internal identifiers in a claimed dump don’t resolve against a company’s own ID space, the data almost certainly didn’t come out of the system that issues those identifiers. What it doesn’t rule out is worth listing precisely, because this is where these disputes usually turn. It doesn’t rule out data from a period predating an identifier migration, for instance. Nor does it rule out data from a third-party processor or partner that maintains its own identifier space. An aggregation is also still possible, where real personal data from multiple sources was assembled and fitted with invented