Strategy·12 min read

B2B Data Provenance Is Now a Buying Criterion

Updately Team·2026-09-19

The two-line court order that sits underneath your prospecting stack

On 17 September 2026, the U.S. District Court for the Northern District of California entered a consent judgment that permanently bars the data broker ProAPIs and its employees from accessing LinkedIn data through fake accounts, bots or other automated technologies — and from selling LinkedIn data at all. Bloomberg Law reported the judgment the same week. If you run outbound, that order is more relevant to your number than most of what came out of conference season, because B2B data provenance — where the contact records in your sequences actually came from — has quietly become a procurement question rather than a legal footnote.

Most sales leaders will read that sentence and assume it does not apply to them. They do not run a scraper. They buy a tool. The tool has a nice UI, a SOC 2 badge, a sales engineer who answers Slack messages, and a per-seat price. Whatever happens two or three hops upstream is somebody else's compliance problem.

That assumption is exactly what the last eighteen months of enforcement has been dismantling. The pattern is now clear enough to plan around: LinkedIn is not trying to win a philosophical argument about the open web. It is removing individual suppliers from the board, permanently, one consent judgment at a time. And when a supplier disappears, the tools sitting on top of it do not send you a warning email. Your match rates just get worse.

What actually happened in the ProAPIs case

LinkedIn filed its complaint in October 2025 in the Northern District of California, naming ProAPIs Inc. and its chief executive (docket 5:25-cv-08393). The allegation was not that ProAPIs had crawled some public pages. It was that the company had created large volumes of fake accounts specifically to extract profile information at scale, then resold access to that data — reporting at the time put subscription pricing as high as $15,000 per month. The Record covered the original filing and iTnews noted the fake-account mechanics that made the case different from a generic scraping dispute.

An agreement in principle surfaced in February 2026. The consent judgment landed on 17 September. The practical outcome: one of the suppliers in the B2B data market is now legally prohibited from being a supplier.

Why this is a pattern, not an incident

Take ProAPIs on its own and it reads as a single bad actor getting caught. Put it in sequence and it reads as a supply chain being cleared.

Proxycurl was a LinkedIn data API that a meaningful number of enrichment products, GTM engineering workflows and internal scripts quietly resold or depended on. LinkedIn sued. The case resolved in LinkedIn's favour, with obligations to stop unlawful access and delete LinkedIn data obtained that way, and Proxycurl shut down the service. Teams that had built enrichment pipelines against that endpoint found out when their jobs started failing.

Before that, hiQ Labs spent years litigating the question of whether scraping public profiles violated the Computer Fraud and Abuse Act. The Ninth Circuit's answer — that accessing public data likely does not trigger the CFAA — is still the line most scraping vendors quote in their sales decks. What they quote less often is how the hiQ matter actually ended: LinkedIn prevailed on breach of contract grounds, and hiQ did not survive it. "Not a federal crime" and "a business you can build on" turned out to be very different claims.

The strategic read is straightforward. Contract law, fake-account and trademark theories, and consent judgments have proven more effective than the CFAA ever was. Each one is narrow, each one is permanent, and each one takes a supplier off the field without setting a precedent anyone else can appeal.

The three-layer supply chain most GTM teams cannot see

The reason a district court order in California can degrade a sequence in Manchester is that B2B contact data moves through at least three layers before it reaches your CRM, and buyers usually only ever meet the third one.

Layer one: the source

Someone has to acquire the underlying record. That happens through some combination of public web crawling, platform APIs under licence, contributory networks where users trade their contact book for access, opt-in forms, partnerships, and — in the cases above — automated collection through accounts that should not exist. From the outside, these produce records that look identical. A name, a title, a company, an email, a LinkedIn URL. Nothing in the row tells you which method produced it.

Layer two: the aggregator

Sources get blended. Aggregators buy from multiple upstream providers, merge and dedupe, run validation, and resell a combined dataset. Blending is genuinely useful — it is how coverage and match rates get to usable levels. It is also where provenance goes to die, because once three sources are merged into one record, the lawful ones and the unlawful ones are indistinguishable downstream.

Layer three: the tool you actually bought

Your sequencer, your enrichment app, your AI SDR, your Clay-style orchestration layer. This is the vendor whose name is on your invoice, whose DPA your legal team reviewed, and whose logo sits in your stack diagram. It may not own a single record. It may be reselling layer two, which blended four of layer one, one of which is now enjoined.

This is why "we use a reputable vendor" is not an answer to a provenance question. Reputation attaches to layer three. Risk originates at layer one.

Four ways provenance risk actually shows up in your number

Provenance sounds abstract until it is a missed quarter. Here is how it surfaces in practice:

  • Silent coverage collapse. A supplier gets enjoined or shuts down. Your vendor backfills from remaining sources. Nobody announces it. Your enrichment match rate drifts from 78% to 61% over a quarter, your sequences start hitting stale titles, and your reply rate decline gets blamed on messaging.
  • Staleness that looks like bad targeting. Prohibited sources cannot be refreshed. Records that came from them freeze at their last-collected state. Job-change and promotion signals — some of the highest-converting triggers in outbound — decay first, because they are exactly the fields that require continuous recollection.
  • Procurement and security review friction. Enterprise buyers, and increasingly mid-market ones, now ask where your data comes from as part of vendor security review. If your own GTM stack cannot answer that, it becomes a deal-slowing question in your customers' process, not just yours.
  • Regulatory exposure that is not about LinkedIn at all. GDPR and its analogues require a lawful basis for processing personal data, and legitimate interest assessments assume you can describe the source. "Our vendor blended it from providers they will not name" is a weak position in a data subject access request, and a worse one in a regulator's inbox.

None of these arrive as a lawsuit. They arrive as a metric that quietly gets worse while everyone debates subject lines.

How the sourcing models actually compare

Not all data acquisition carries the same risk, and it is worth being precise rather than treating everything as either fine or forbidden.

Sourcing modelHow it worksDurabilityPractical risk
Licensed platform APIVendor pays for sanctioned access under contractHigh, but subject to pricing and terms changesLow legal risk; high concentration risk if terms shift
First-party signalsData generated by your own accounts and interactionsVery high — you own the relationshipMinimal, provided consent and disclosure are handled
Opt-in and contributory networksUsers exchange data for accessModerateConsent quality varies enormously between providers
Public web crawlingAutomated collection of genuinely public pagesModerateCFAA risk is low; terms-of-service and contract risk is real
Fake-account automated collectionSynthetic accounts extract gated data at scaleNoneThe model the ProAPIs and Proxycurl judgments removed
Undisclosed blended resaleAggregated from unnamed upstream providersUnknown by definitionYou inherit the worst source in the blend

The row that matters most for 2027 planning is the last one. It is not that blended resale is inherently improper. It is that a vendor who will not tell you what went into the blend has transferred an unquantifiable risk onto your balance sheet and called it a feature.

The vendor provenance audit: nine questions to send this week

You do not need a legal team to run this. Send these to every data and enrichment vendor in your stack, in writing, and keep the replies.

  1. Which upstream providers contribute to the records you sell us? Name them.
  2. For each, what is the acquisition method — licensed API, opt-in, contributory network, public crawl, or other?
  3. Have any of your upstream providers been the subject of litigation, an injunction, or a consent judgment in the last 24 months?
  4. What is your contractual obligation to notify us if a source is removed or a dataset must be deleted?
  5. What was our match rate by field over the last four quarters? Has it moved?
  6. Do you collect any data through accounts created for collection purposes, directly or through a supplier?
  7. What is your documented lawful basis for processing EU and UK personal data, and can you produce the legitimate interest assessment?
  8. If a source disappeared tomorrow, what percentage of records in our contracted segments would go stale, and how quickly?
  9. Will you commit to the answers above in the contract, not just in email?

Two things to watch for in the responses. First, any vendor that treats question one as commercially sensitive has effectively answered it. Second, question five is the honest test, because match-rate history is a fact they already have. A vendor who will not share their own trend line is telling you the trend line is bad.

The structural answer: stop renting the relationship

Auditing vendors is necessary hygiene. It is not a strategy. The strategy is to shift the weight of your pipeline generation from data you rent toward signals you generate, observe, or earn — because those do not have an upstream supplier who can be enjoined.

Signals you generate

Every piece of content your team publishes creates a measurable interaction surface. Who viewed your profile this week. Who engaged with the post your VP of Sales wrote about pricing models. Who attended the webinar and stayed past the twelve-minute mark. These are not purchased records; they are behaviour directed at you. They carry consent context, they are current by definition, and no third party can take them away.

They are also, empirically, better leads. The gap between a cold record and someone who engaged with your content last Tuesday is not a small lift in reply rate — it is a different motion entirely. This is the core of what signal-based outbound means, and it is the reason Updately is built around capturing profile views, post engagers, competitor mentions, hiring signals and pain-point posts first, then enriching and scoring against ICP, rather than starting from a purchased list and hoping.

Signals you observe with permission

Public, timestamped, intent-rich behaviour exists in enormous volume and does not require anyone to run a fake account. Someone asking for tool recommendations on Reddit. A company posting three roles that only make sense if a specific initiative is funded. A funding announcement. A leadership change. An executive publicly complaining about the exact problem you solve. The value here is not the identity — it is the timing, and timing is the one thing purchased lists structurally cannot give you.

Signals you earn

Customers, former colleagues, investors, and community members produce warm paths that no database can replicate. Most teams underuse these because they are harder to put in a spreadsheet, not because they convert worse. They convert dramatically better.

The point is not that you should never buy data. Enrichment is still how you get from a signal to a complete record you can act on. The point is where the trigger comes from. If your pipeline starts with a purchased list, your pipeline inherits every risk in that list's supply chain. If it starts with an observed signal and uses purchased data only to fill in details, a supplier disappearing becomes an inconvenience rather than an outage.

What to do in the next 30 and 90 days

This week. Send the nine questions. Pull your enrichment match rate by field for the last four quarters and look for the drift. Identify which of your sequences depend on a single data source for their trigger.

Next 30 days. Add a provenance clause to every data vendor renewal: named upstream sources, notification obligations on source removal, and deletion cooperation. Stand up at least one signal source you own — profile views, post engagement, or content interaction — and route it into a real sequence so you have a working comparison.

Next 90 days. Run the comparison honestly. Measure reply rate, meeting rate and cost per qualified opportunity for purchased-list-triggered outbound against signal-triggered outbound, using the same reps and the same messaging. Most teams that run this test find the difference large enough that the provenance question answers itself. Then rebalance the budget accordingly, and make source diversity an explicit requirement rather than an accident.

Takeaways

  • A consent judgment entered on 17 September 2026 permanently barred a LinkedIn data broker from accessing or reselling platform data — the latest in a sequence that has already removed Proxycurl and hiQ from the market.
  • LinkedIn is winning on contract, fake-account and trademark theories rather than the CFAA, which means each outcome is narrow, permanent, and not appealable by anyone else.
  • B2B data provenance is now a procurement question. Most GTM teams buy at layer three of a three-layer supply chain and cannot name layer one.
  • Provenance failures do not announce themselves. They appear as declining match rates, stale job-change data, slower security reviews, and weak answers to regulators.
  • Run the nine-question vendor audit this week, and treat a refusal to name upstream sources as an answer.
  • The durable fix is to trigger outbound from signals you generate, observe, or earn, and use purchased data for enrichment rather than for the trigger itself.

The teams that will be fine in 2027 are not the ones with the biggest database. They are the ones who can say, in one sentence, where every record in their pipeline came from — and whose best leads never needed a database in the first place.