A property valuation can be accurate and indefensible at the same time. Since October 2025 that distinction has had regulators attached to it.

Executive brief60-second version
Business problem
A property valuation can be accurate and indefensible at the same time. Since October 2025 that distinction has had regulators attached to it, and the rule deliberately does not say how to satisfy it.
Why it matters
Six federal agencies now require quality controls across five standards for automated valuation models in covered mortgage decisions — confidence, protection against manipulation, conflicts of interest, random sample testing, and nondiscrimination. A portfolio accuracy metric is not a per-property confidence, and aggregate accuracy hides local error.
Architecture response
Treat the valuation as a record rather than a number: point-in-time fixtures so a backtest is not run against today’s table, retained comparable sets, lineage that makes the manipulation and conflict questions answerable, and error reported by segment rather than in aggregate.
What Databright Cloud Solutions does
We design point-in-time valuation fixtures, calibration and segment-level error reporting, comparable-set retention, field-level provenance, and the valuation record and review workflow that makes an estimate reproducible.

The rule implements a Dodd-Frank mandate and took effect on 1 October 2025. It reaches mortgage originators and secondary market issuers using an AVM to determine the collateral value of a mortgage secured by a consumer’s principal dwelling.1 Notably, it names five outcomes and then declines to prescribe structure: the agencies stated plainly that the rule “does not set specific requirements for how” those policies and control systems should be built, so that institutions can scale them to their size and risk.2

That flexibility is sensible regulation. It is also a transfer of responsibility.

The rule names what must be true. Deciding what would make it true is a data problem.

A portfolio metric is not a per-property confidence

Ask a valuation provider about accuracy and you will usually receive a median absolute error and a hit rate — the share of estimates within ten percent of the eventual sale price. Both are real, useful numbers. Both describe a population.

Neither says anything about the house in front of you. A model with a four percent median error can be badly wrong on a specific property, and the standard concerns confidence in the estimate being produced, not the average estimate across a book.

The engineering form of that requirement is calibration. If a model emits a ninety percent interval, the eventual sale price should land inside that interval about ninety percent of the time. That is measurable, and it is frequently not measured.

What a confidence number has to survive
EmitValue plus interval
ObserveThe eventual sale
CoverageDoes 90% mean 90%?
SegmentType · band · geography

Segmentation is where this gets uncomfortable. An interval that is well calibrated nationally is routinely overconfident on unusual property types, at the top and bottom of the price range, and anywhere transaction density is thin. A single global coverage number hides exactly the cases where a valuation is most likely to be challenged.

One thing worth checking early: whether the confidence score you are relying on is a confidence at all. Some are data-completeness proxies wearing a probability’s clothing — higher when more fields were populated, without any claim about error. That is a useful signal and it is not calibration.

The model that matters is comparable selection

Attention naturally goes to the estimator. Most of the error arrives before it.

Which sales were eligible, whether they are genuinely the same property type, whether an arm’s-length filter excluded the transfers it should have, whether “living area” means the same measurement across every comparable — these decisions bound the achievable accuracy before any coefficient is applied. Comparable selection rests on resolving whether two records describe the same parcel or unit,6 and on field definitions that are standardized in name more reliably than in meaning.7

It follows that when two valuations of the same property disagree materially, the productive question is rarely about the algorithms. It is which comparables each one used.8

A practical consequence

Store the comparable set alongside the estimate, with the identifiers and the characteristics as they stood at valuation time. An estimate whose comp set cannot be reconstructed cannot meaningfully be reviewed — by you, by an auditor, or by anyone contesting the number.

A backtest against today’s table is not a backtest

The fourth standard calls for random sample testing and reviews. The obvious implementation is to sample past valuations and compare them against subsequent sales.

The obvious implementation is also where the result quietly becomes meaningless. Query today’s warehouse and the sample inherits everything that arrived afterwards: corrections applied to records after the valuation date, sales recorded late, characteristics revised by a permit that closed months later. The model is being graded with information it did not have, and it will look better than it was.3

Testing that measures the model rather than hindsight
SamplePast valuations
ReconstructInputs as of that date
FreezeImmutable test set
CompareAgainst the outcome

Freezing matters as much as reconstructing. If the inputs behind a test can drift between runs, a quarter-over-quarter comparison cannot separate a model change from a data change — and the review record shows a trend that may be neither.

Manipulation and conflicts are lineage questions

Two of the five standards — protecting against data manipulation and avoiding conflicts of interest — read like policy commitments. Inside a data platform they are both answerable only with lineage.

The operative question is whether a party with an interest in the outcome can influence an input. Owner-supplied characteristics, agent-entered remarks, a record corrected shortly before a valuation was ordered: none of these is inherently improper, and all of them are worth being able to see. Answering that requires field-level provenance — which system supplied each value, when it was observed, and what transformed it on the way in.5

There is a second-order version of this that is easy to miss. If you consume several vendor valuations and select or blend among them, the selection rule is itself a model, and it belongs under the same governance as any other. Taking the highest of three estimates is not a valuation method. It is a policy, and an adverse one to have to explain later.

Aggregate accuracy hides local error

The fifth standard concerns compliance with applicable nondiscrimination laws. The measurement question underneath it is narrower and entirely tractable: error is not evenly distributed, and an aggregate figure is built to conceal that.

A national median error is perfectly compatible with systematically larger error in particular places — and model error tends to concentrate where transaction density is low, where property characteristics are recorded less completely, and where housing stock is less uniform. Those conditions are not randomly distributed across geographies.

So error has to be reported by segment, geography included, and reviewed as a distribution rather than a headline. This is the same discipline any model making property claims requires: compare populations of outputs, not individual answers.4 Whether a given result raises a legal issue is a question for counsel; whether you can see the result at all is a question for your data platform, and it is the one you can settle in advance.

What a valuation record has to carry

Every standard above converges on the same requirement: a valuation must be reproducible long after the moment it was produced. That is a schema decision more than a modeling one.

The minimum a valuation should carry

Enough to reconstruct the estimate, the evidence, and the reasoning, without access to the model that produced it.

value: 612000
as_of: 2026-09-15
interval: [571000, 656000]  coverage: 0.90
model: avm-r4.2  ruleset: comps-2026.03
input_snapshot: snap-8f21c4
comparables: [P-00184, P-00921, P-01337]
provenance: per-field, retained
override: none

The two fields most often missing are the snapshot pointer and the override. Without a frozen input snapshot the estimate cannot be recomputed. Without an explicit override field, human adjustments become invisible — and an adjustment nobody recorded is indistinguishable from a model output when someone asks a year later who decided the number.

Defensible beats accurate

An accurate valuation you cannot explain is a liability. A slightly less accurate one that carries its interval, its inputs, its comparables, and its provenance is an asset, because it can be reviewed, challenged, corrected, and improved.

The rule’s flexibility is the part worth sitting with. Because it does not define what a high level of confidence is, that definition becomes yours to write down, defend, and test against — which is a better position than it sounds, provided the writing down happens before the testing does.

The principle

Treat confidence as a claim to be calibrated, not a score to be displayed. Reconstruct the past rather than filtering the present. Keep provenance on every input, the comparable set with every estimate, and human overrides on the record. Measure error by segment, including geography. Then a valuation is not merely a number — it is a number you can still account for a year later.

Sources and industry references

  1. CFPB — Quality Control Standards for Automated Valuation Models: interagency final rule (CFPB, OCC, Federal Reserve, FDIC, NCUA, FHFA), effective 1 October 2025
  2. OCC Bulletin 2024-17 — Quality Control Standards for Automated Valuation Models: final rule, and its flexibility on implementation
  3. Databright Cloud Solutions — The Temporal Truth Problem: what was true, when it became true, and when we knew it
  4. Databright Cloud Solutions — Right Property, Right Time, Right Evidence: evaluating real estate AI
  5. Databright Cloud Solutions — The Source Provenance Problem: knowing where every property value came from
  6. Databright Cloud Solutions — The Property Identity Problem: why APNs, parcel numbers, addresses, and MLS IDs are not enough
  7. Databright Cloud Solutions — The Semantic Consistency Problem: why matching field names do not mean matching definitions
  8. Databright Cloud Solutions — The Property Data API Paradox: more sources, less agreement

This article provides data-architecture and engineering perspectives on valuation systems. It is not legal, compliance, appraisal, lending, investment, or regulatory advice, and it is not a statement of what any particular institution must do to satisfy the rule. Regulatory scope, applicability, and expectations depend on the institution, the transaction, and facts this article does not address; consult qualified counsel and your regulators.