Ask an engineer where a property fact came from and you get a pipeline. Ask a lawyer and you get three separate questions, each with a different answer, none of which is settled by the fact that the data is sitting in your warehouse.
- Business problem
- Contract, copyright, and privacy law govern the same property record at the same time, and they do not resolve into a single permission. The record sitting in your warehouse settles none of them.
- Why it matters
- A use can be contractually permitted and still unlawful, or lawful and still a license breach. MLS data reaches you by contract rather than by right, “public record” does not mean unrestricted, copyright covers the listing photograph but not the square footage, and some individuals hold rights over their own address.
- Architecture response
- Model all three regimes separately and per field: versioned license terms, jurisdiction-aware policy, AI training treated as its own permission, and suppression that reaches indexes and training sets rather than only the primary table.
- What Databright Cloud Solutions does
- We design field-level rights models, jurisdiction-aware policy, versioned license terms, global suppression that reaches indexes and training sets, and the provenance needed to answer a question about a decision made months ago.
This is an engineering perspective, not legal advice, and it is deliberately not a comprehensive survey. The United States has no single property-data law: rules differ by state, by county, by MLS, and by contract, and they change. What follows is a handful of illustrative examples chosen to show the shapes of constraint a platform has to model — not a compliance checklist, and not a substitute for counsel who knows your jurisdictions and your agreements.
The previous article in this series argued that usage rights should be machine-readable policy rather than prose in a contract.1 That article was about mechanism. This one is about what the mechanism has to encode — and why the answer is rarely a single flag on a table.
Three regimes stack on the same record
Consider one row: a single-family home, listed last spring, sold in June, currently owned by an LLC. That row may simultaneously be:
The listing came through an MLS data feed governed by a participation agreement, IDX or VOW rules, and a downstream vendor contract — each with its own permitted uses, display requirements, and retention terms.
The photographs and, in some cases, the remarks are original expression owned by a photographer, agent, or brokerage — while the underlying facts are not protected at all.
The deed, parcel, and assessment data came from a county office that is required to make it available — subject to state rules on bulk access, fees, redaction, and who may be identified.
The owner name and mailing address may be personal information under one or more state privacy statutes, and the individual may hold statutory rights over it regardless of where you obtained it.
These regimes do not resolve into a single permission. They intersect. A platform that stores one source column and one license flag has already lost the ability to answer the question correctly.
MLS data reaches you by contract, not by right
The most common engineering mistake is treating an MLS feed like an open data source with an access key. It is closer to the opposite: a tightly conditioned license, where the conditions are the point.
Typical MLS and downstream agreements speak to display context, attribution, refresh and takedown obligations, retention of sold or expired records, redistribution to third parties, derivative products, and increasingly whether the content may be used to train models. Those terms differ between MLSs, and a national platform may hold hundreds of agreements that do not agree with one another.
This is precisely why the semantic-consistency problem and the rights problem are the same shape: the terms are written in prose, for humans, per market. Nothing about the feed itself tells your runtime that market A permits a derived valuation product and market B does not.
Rule changes move data, not just behavior
Industry rule changes are often read as agent-practice news. They are also schema events.
Following the 2024 settlement of the commission litigation, NAR’s practice changes took effect on August 17, 2024. Among them: offers of compensation may no longer be communicated through the MLS, and MLS participants working with buyers must enter into written agreements before touring a home.2
For a data platform, the first of those is a field disappearing. Historical rows retain values that current rows cannot have. Any metric computed across that boundary — and any model trained on it — is comparing two different regimes. This is the temporal-truth problem wearing a legal hat: the schema changed because the rules changed, on a specific date, and the platform needs to know that date.
“This field is null after August 2024” is not a data-quality defect to be cleaned. It is the correct representation of a legal change, and erasing it corrupts every longitudinal analysis that crosses it.
Availability is not permission
In March 2025 NAR introduced a policy allowing sellers to delay public marketing of a listing. Under “delayed marketing exempt” status, a listing remains visible to MLS participants and subscribers while being withheld from IDX feeds and syndication for a period each MLS sets, with a signed seller disclosure required.3
The engineering consequence is direct: a record can be present in your ingest and still be prohibited from your public surface. If the only thing your pipeline models is “we have it,” it will publish something a seller specifically opted out of publishing. Display eligibility has to be a first-class attribute traveling with the record, not an assumption derived from its presence.
Copyright protects the photograph, not the square footage
Copyright covers original expression, not facts. A bedroom count, a lot size, a sale price, and a parcel number are facts; no one holds copyright in them. Photographs, video, floor-plan drawings, and sufficiently original written remarks are a different matter — they are creative works with an author and an owner.
Ownership of listing photographs is frequently misunderstood in exactly the direction that creates liability. NAR’s own guidance to members addresses who owns property photos and cautions that photographers commonly retain copyright, licensing limited use to the agent or brokerage rather than transferring ownership.4 An MLS may hold rights in the compilation without holding rights in the underlying images.
Beds, baths, area, price, parcel identifiers, dates. Not protected by copyright, though a license may still restrict what you do with them.
Photos, video, tours, floor plans, original remarks. Protected, with an owner who may not be the MLS or the brokerage.
The selection and arrangement of the database may attract thin protection distinct from the underlying items.
Enhancements, virtual staging, and model outputs derived from a protected image raise their own questions about who holds what.
The practical implication for a platform is that media needs a different rights model from structured fields. The same listing can carry facts you may freely analyze and images you may not retain past the listing’s active period.
“Public record” does not mean unrestricted
Deeds, mortgages, assessments, and parcel maps are created by government offices under statutory duties to make them available. That availability is real, and it is also narrower than the phrase suggests.
Access terms vary by state and often by county: some jurisdictions distinguish inspection of individual records from bulk extraction, charge differently for bulk products, impose conditions on commercial redistribution, or require redaction of specified identifiers. Two counties in the same state can present materially different terms. A national platform ingesting all of them inherits all of those differences at once.
State privacy law: the carve-out is narrower than it looks
By 2026, roughly twenty US states have comprehensive consumer privacy statutes in effect, with additional states enacted and phasing in.5 There is still no general federal analogue, so a national property platform is operating under a patchwork.
Most of these statutes exclude “publicly available information” from the definition of personal information, and most definitions include information lawfully made available from government records. California’s definition is the most-cited example.67 Teams often read that as settling the question for property data. Three reasons to be careful:
States word the carve-out differently, and what qualifies in one state may not in another. A single global “public record, therefore exempt” flag is a modeling error, not a compliance position.
A public record combined with purchased consumer attributes, inferred household composition, or model-derived scores is no longer only the public record. The carve-out attaches to information, not to the file it lives in.
Sector rules, contractual terms, marketing and telephony restrictions, and state-specific statutes can apply to information that is not “personal information” under a given privacy act.
Some individuals hold rights over their own address
The sharpest illustration that “public record” is not the end of the analysis is the class of statutes protecting the home addresses of specific officials.
New Jersey’s Daniel’s Law lets covered individuals — judges, prosecutors, law-enforcement and corrections officers, and their immediate family — request in writing that a company not disclose their home address or unpublished phone number, with statutory damages of $1,000 per violation and a short window to comply. The resulting litigation has been substantial: a single assignee entity has brought well over a hundred suits on behalf of roughly nineteen thousand covered individuals.8
Several states have enacted address-confidentiality protections of varying scope, and the details differ meaningfully. The architectural lesson generalises even where the specific statute does not:
A platform must be able to suppress a specific individual’s address across every derived product, index, cache, export, and model input — on request, within days, and provably. If suppression is a filter applied at one display surface, it is not suppression.
That is a hard engineering requirement, and it is the one most likely to be discovered late. It reaches search indexes, materialised aggregates, cached tiles, partner exports, embeddings, and training corpora. It is also the clearest case where provenance and identity resolution stop being data-quality concerns and become the mechanism by which a legal obligation is actually met.
Scraping: the CFAA narrowed, everything else did not
Teams sometimes reason that if a page is publicly reachable, collecting it is permitted. The case law is more specific than that, and considerably less comforting.
In Van Buren v. United States (2021) the Supreme Court read the Computer Fraud and Abuse Act’s “exceeds authorized access” narrowly, rejecting a reading that would turn terms-of-service violations into federal crimes.9 The Ninth Circuit, on remand in hiQ Labs v. LinkedIn, reaffirmed that scraping publicly available data likely does not constitute access “without authorization” under the CFAA.10
What that pair of decisions did not do is make scraping generally lawful. The CFAA is one statute among many. Breach of contract, copyright infringement, trespass to chattels, misappropriation, and state statutory claims were untouched — and in the same litigation, LinkedIn ultimately prevailed on a contract theory. For property data specifically, the material terms usually live in an MLS or vendor agreement your organization has actually signed, which is a contract question long before it is a CFAA question.
Fair housing constrains use, not just access
Everything above concerns whether you may hold and move data. A separate body of law governs what you may do with it.
The Fair Housing Act prohibits discrimination in housing-related transactions on the basis of protected characteristics.11 For a platform, the exposure is rarely an explicit rule keyed to a protected class. It is a model that reproduces a protected characteristic through correlated features, or a targeting product that shapes who is shown which properties, or an automated valuation whose errors fall unevenly across neighbourhoods.
This is where provenance stops being a governance nicety. If a scoring or targeting decision is challenged, the question is which features contributed, where those features came from, and what the system knew at the time — the reconstruction problem from the temporal-truth and provenance articles, asked by someone with subpoena power.
AI training is its own permission
The right to display a listing, the right to retain it, and the right to train a model on it are three different permissions, and older agreements frequently addressed only the first.
The industry is actively working on making this machine-readable rather than contractual prose, which the previous article covered in detail.1 Until those mechanisms are widely implemented, the practical position is that silence in a 2019 agreement is not consent for 2026 model training, and a platform that cannot say which corpora a model was trained on cannot answer the question at all.
What this means for the platform
The legal complexity is not going to resolve into a simpler rule. The engineering response is to stop trying to compress it into one and to model the dimensions that actually vary.
One license flag per dataset cannot express “facts analysable, photos not retainable past close.”
State and county are inputs to the permission decision, not just address components.
Agreements and industry rules change. A permission decision must be reproducible as of a past date.
Individual-level takedown must reach indexes, caches, exports, aggregates, and training sets — with evidence it did.
They are distinct permissions and are routinely granted separately.
Provenance is how you answer a regulator, a plaintiff, or a partner audit about a decision made months ago.
The unglamorous conclusion
There is no single answer to “may we use this property record.” There is a license, a copyright position, a public-records regime, a state privacy statute, possibly an individual’s statutory right over their own address, and a separate body of law about what the output may be used to decide.
The organizations that handle this well are not the ones with the most permissive contracts. They are the ones whose systems can state, for any given field, where it came from, which agreement governed it, what that agreement permitted as of a specific date, and what was actually done with it.
Legal exposure in property data is rarely created by a decision to break a rule. It is created by an architecture that cannot tell which rule applied.
Sources and industry references
- Databright Cloud Solutions — The Usage Rights Problem: Making Real Estate Data Governance Machine-Readable for APIs, Analytics, and AI
- National Association of REALTORS® — Settlement FAQs: practice changes effective August 17, 2024, including removal of compensation offers from the MLS and written buyer agreements
- National Association of REALTORS® — Multiple Listing Options for Sellers: delayed marketing exempt listings (March 2025)
- National Association of REALTORS® — Who Owns Your Property Photos? Copyright ownership and licensing of listing photography
- MultiState — Comprehensive state consumer privacy laws in effect and taking effect in 2026
- California Attorney General — California Consumer Privacy Act overview
- California Civil Code § 1798.140 — definitions, including “publicly available” information
- Ogletree Deakins — New Jersey’s Daniel’s Law: takedown obligations, statutory damages, and the resulting data-broker litigation
- Van Buren v. United States, 593 U.S. ___ (2021) — narrowing “exceeds authorized access” under the Computer Fraud and Abuse Act
- hiQ Labs, Inc. v. LinkedIn Corp., No. 17-16783 (9th Cir. Apr. 18, 2022) — scraping publicly available data and the CFAA
- U.S. Department of Justice, Civil Rights Division — The Fair Housing Act
This article is not legal advice, and it is not a comprehensive or authoritative statement of the law. It is a data-architecture and technology perspective written for engineering and product teams. The examples above were selected to illustrate categories of constraint; they are a small and non-exhaustive sample, and many states, counties, MLSs, and statutes that would matter to a given platform are not mentioned at all. Laws, industry rules, and contractual terms differ by jurisdiction and change over time, and litigation described here may be ongoing or superseded. Nothing here creates an attorney-client relationship. Consult qualified counsel familiar with your jurisdictions, your agreements, and your intended uses before making compliance, licensing, or product decisions.