We have argued elsewhere that pointing a model at a warehouse produces fluent answers rather than correct ones.1 This article assumes you accept that and want to build the thing that fixes it. It is about construction: what to define, in what order, and how the model is allowed to ask.
- Business problem
- Leaders want to ask questions of company data in plain language. The pilot that wires a model to the warehouse demos well and cannot be trusted with a number anyone will act on, because nothing in it decides what “revenue” means, which join is valid, or who is allowed to see the result.
- Why it matters
- The failure is not refusal, it is confident precision. A model given raw tables will produce a number with two decimal places that no finance team recognizes, and there is nothing in the answer that reveals which of several plausible definitions it used. Consistency across a business cannot be prompted into existence; it has to be defined once and enforced somewhere.
- Architecture response
- Do not let the model write SQL. Publish a catalog of named metrics and dimensions, let the model discover it and request entries by name in a structured query, and have the layer validate that request, apply access policy deterministically, compile the SQL itself, and return the result together with the definitions used. Then test the layer with a golden question set, because the compiled query is deterministic and therefore assertable.
- What Databright Cloud Solutions does
- We model the entities, dimensions, and metrics with the people who own the definitions, build the governed query surface the model talks to, wire access policy into the layer rather than the prompt, and stand up the evaluation set that tells you the layer is right before a business decision depends on it.
The interface decision that determines everything else
There are two ways to connect a language model to analytical data, and choosing between them settles most of the subsequent architecture.
In the first, the model receives schema information and generates SQL, which is executed. This is the default pattern in demos because it requires almost no preparation. Its failure mode is structural: the model is being asked to invent business definitions at query time, from names alone, with no way to signal that it has guessed. As one semantic-layer vendor puts it directly, agents writing SQL against a warehouse end up with inconsistent metrics and ungoverned access — numbers that do not match how the business defines them.2
In the second, the model queries a semantic layer. It asks what is available, receives a catalog of governed measures and dimensions with descriptions rather than raw table definitions, and then requests them by name. The pattern that has consolidated around this in 2026 uses the Model Context Protocol as the discovery and query surface: the agent lists the available measures and dimensions, then asks for them by name rather than writing SQL, against a typed API that returns self-correcting validation errors.3
The second is more work up front and it is the one that survives contact with a finance team. Everything below assumes it.
The four objects you have to define
A semantic layer is not a glossary. It is a queryable model with a small number of object types, and getting their boundaries right is most of the work. The vocabulary below follows the widely adopted modeling approach and maps cleanly onto the open specification discussed later.
The starting points, corresponding to tables in your transformation project, carrying the metadata — table name, primary keys — that lets the graph be navigated correctly.4 One table can yield several semantic models when it plays several roles.
The join keys, typed as primary or foreign — the traversal paths between models.4 Typing them is not bookkeeping; it is what lets the layer work out which joins are legitimate.
The ways you group and slice, either time or categorical. Without them a metric is simply a number for all time.4 A dimension attaches to the primary entity of its model and is not freely joinable to metrics elsewhere.
Functions combining simple column references, constraints, or other metrics into quantitative indicators.4 Simple, ratio, cumulative, derived, and conversion types cover the large majority of real business questions.
The discipline that pays off is refusing to define a metric twice. If net_revenue exists, a second metric that is net revenue with one region excluded is not a new metric — it is the existing metric with a filter. Teams that skip this end up with forty metric names covering nine concepts, and a model choosing among them by name similarity. That is the original problem with extra steps.
Definitions belong in version control as declarative files, reviewed like code, so that everyone on the data and business teams can see and approve them as the single source of information.4 The review is the point. A metric definition is a business agreement that happens to be machine-readable, and the argument about what counts as an active customer is far better had in a pull request than discovered in a board pack.
Join safety is a modeling problem, not a prompting problem
The most damaging analytical errors are not syntax errors. They are joins that execute perfectly and return numbers that are too high.
Fan-out is the common case: join an order header to order lines and sum the order total, and every order is counted once per line. The query succeeds. Revenue is inflated by the average basket size. A chasm join — two unrelated one-to-many branches joined through a shared parent — produces a similar multiplication with an even less obvious cause. No model, however capable, reliably detects these from column names, because the information needed is cardinality and grain, and that is not in the names.
This is exactly what entity typing buys. A semantic layer that captures the type of each identifier can navigate to appropriate joins and avoid the construction of fan-out and chasm joins while generating legible SQL.4 The protection is structural: the wrong join is not something the model declines to write, it is something the query interface cannot express.
Make incorrect questions unrepresentable rather than discouraged. Anything you are relying on the prompt to prevent will eventually happen, because prompts are advisory and interfaces are not. Every constraint you can move from instruction into type is a class of wrong answer permanently retired.
Time grain is where most wrong answers come from
Ask a business for revenue in September and the honest answer is a question: revenue recognized in September, invoiced in September, or booked from orders placed in September? Those are three different numbers, all defensible, all called revenue.
A semantic layer has to make the choice explicit and then make it hard to get wrong. In practice that means time dimensions carry an intended granularity, metrics declare which date they are measured against, and the query interface requires a grain rather than defaulting to one silently. A default is fine. An undeclared default is how two dashboards disagree for a year.
Two further cases deserve first-class treatment because business users assume them and models cannot infer them. Period-over-period comparison needs to be a defined capability, not something assembled per query, or you will get comparisons against calendar periods where the business means fiscal ones. And late-arriving data means today’s figure for last month can change after the fact — which is a property of your pipeline that the answer should disclose rather than hide. Separating when something became true from when you learned it is its own discipline, and worth resolving before a model starts reporting on it.5
The model reads your names, so write them for a stranger
Here is the part teams consistently underinvest in. What the model sees is not your warehouse. It is the catalog of governed measures and dimensions with their descriptions, not raw table definitions.3 That catalog is the prompt. Its quality determines selection accuracy more than any prompt engineering downstream.
Write for a competent analyst on their first day who will not ask a follow-up question. State what the metric counts, what it excludes, and the unit. “Revenue” is not a description; “net revenue recognized, excluding intercompany, in reporting currency” is.
Where two metrics are genuinely similar, each description should say when to prefer the other. This is the single highest-yield edit available, because near-neighbors are where selection errors cluster.
The business says churn, logo churn, and attrition for one thing and three things depending on who is speaking. Record the aliases against the metric so resolution is a lookup rather than an inference.
Publishing four hundred metrics because they exist degrades selection. Publish the ones that answer real questions and expand on evidence of need.
The shape of the query the model emits
The model should produce a structured object, not a string of SQL. Something with this shape, validated before anything reaches the warehouse:
Three properties make this work. The request is closed — every name must exist in the catalog, so a hallucinated metric is a validation error rather than a query. It is inspectable — a human or a test can read the request and judge whether it matches the question asked, which is far easier than reviewing generated SQL. And it is compilable — the layer generates the SQL, so join paths, grain handling, and dialect quirks are decided by code that was reviewed once rather than by a model on every request.
Validation errors should be returned to the model to retry against, not surfaced to the user. A typed surface with self-correcting validation errors lets the agent converge on a valid request without ever writing SQL.3 An unknown dimension name is a recoverable event; give the model the list of valid options and it will usually correct on the next turn. Cap the retries, and treat exhaustion as a refusal rather than a fallback to guessing.
Follow-up questions are query edits, not new questions
Nobody asks one question. The second turn in every real session is a refinement — “now split that by plan tier”, “same thing for last quarter”, “excluding trials” — and how the architecture handles it separates a usable system from a demo.
The naive implementation re-derives a fresh query from the conversation text each turn. It drifts. The grain silently changes between turn two and turn four, a filter established early is dropped, and the user has no way to see it happen because each answer looks reasonable on its own. Prose is a poor carrier of accumulated state.
The better design keeps the last validated semantic query as the state and treats each follow-up as a delta against it: add a dimension, change the grain, tighten a filter, swap the metric. Because that state is a small structured object, the transformation is inspectable and reversible in a way a conversation transcript is not.
Two practices make this hold up. Echo the resulting query, or at least what changed, alongside each answer — “added plan tier, kept the monthly grain and the fiscal-year filter” — so drift becomes visible at the turn it occurs rather than three answers later. And let the user reset explicitly, because a long session accumulates filters nobody remembers agreeing to, and “start again from net revenue” should be one instruction rather than a new session.
This is also what makes a session auditable. When someone asks why the number moved between two turns, the answer is a diff between two structured objects, not an archaeology exercise across a chat log.
Access policy belongs in the layer, applied deterministically
The temptation is to handle permissions in the prompt: tell the model which regions this user may see. This is not a control. It is a request, made to a component whose whole purpose is generating plausible text, on behalf of a user who may be actively trying to get around it.
Policy has to sit where it cannot be talked out of. The property worth insisting on is that every query passes through the semantic layer runtime, where it is validated against the data model and has access policies applied deterministically before reaching the warehouse.2 The model's identity is not what authorizes the query; the end user's is, propagated through the layer and enforced there.
Two consequences follow. Row-level and column-level restrictions must be expressible in the layer rather than bolted on, because a metric filtered per user is a different compiled query, not a post-filtered result set — filtering after aggregation gives the wrong denominator on every ratio. And the audit record has to capture the resolved query plus the identity it ran under, since "the AI provided this figure" is not an answer to a regulator. Where the data carries its own handling constraints, those constraints travel with the purpose of the request, not with the credentials of the service account.6
Design the refusal path before the happy path
Most questions a business asks of a semantic layer do not map cleanly onto it. The system’s behavior in that moment determines whether people trust it.
The failure to avoid is silent approximation: a question about “customer profitability” answered with gross margin because that was the closest available metric, presented with no indication that a substitution occurred. The answer is wrong in a way that is invisible and will be repeated.
Nothing in the catalog covers it. Say so, name the closest available metrics, and let the human choose. A stated gap is a feature request; a silent substitution is a defect.
Two or more metrics fit. Ask one clarifying question rather than picking. The cost of a question is a few seconds; the cost of the wrong pick is a decision.
The metric exists and this user may not see it at that grain. Refuse on the same terms a BI tool would, and never explain the restriction in a way that leaks the shape of the data.
A valid request that scans far more than expected. Return an estimate and require confirmation. Cost control is part of the contract, not an afterthought.
Return the query, not just the number
An answer that cannot be checked will eventually be disbelieved, usually at the worst moment. Every response from the layer should carry the evidence of how it was produced: the metrics used with their definitions, the filters applied, the grain, the time range, the row count, and a note where data is provisional.
This matters for a reason beyond audit. It changes what happens when the number looks wrong. With provenance, a user reads the definitions and says “that excludes intercompany, which is why it does not match my report” — a thirty-second resolution and a small increase in trust. Without it, the conversation is about whether the AI can be relied on at all, and that conversation does not end well for the AI.
Define it once, portably
A semantic layer is a durable asset. The model in front of it is not — models carry retirement dates and shifting interfaces on someone else’s schedule.7 That asymmetry argues for expressing your semantics in a form that does not belong to any one vendor.
This has moved quickly. The Open Semantic Interchange published the first version of its specification in January 2026 — a vendor-neutral, extensible model for representing semantic layer constructs such as data sets, metrics, dimensions, relationships, and contexts, under an Apache 2 license, with a working group spanning dozens of platform, BI, and catalog vendors, and explicit intent that models be consistently interpreted across tools, platforms, and agentic applications.8 As of September 2026 the project has been accepted into the Apache Incubator as Apache Ossie, with a declarative standard for defining metrics, dimensions, and joins so that every tool and agent works from the same source of truth, plus a typed query surface for agents.3
Treat this the way you would any incubating standard: worth designing toward, not worth betting delivery on. The practical move is to keep your definitions declarative, in version control, and free of logic that only one query engine understands. Portability then becomes a translation problem rather than a remodeling project — which is the difference between adopting a standard later and rebuilding for it.
Test the layer, not just the model
Here is the structural advantage of this architecture, and the reason to prefer it even setting governance aside: the thing the model produces is deterministic and therefore testable.
Build a golden set of business questions in the words people actually use, each paired with the semantic query that correctly answers it. Then assert on the emitted query rather than on the prose. Did the model select net_revenue or gross_revenue? Month or quarter? Did it apply the fiscal-year filter? These are exact comparisons against a small structured object — cheap to run on every change, and immune to the phrasing variation that makes output-text assertions so brittle.
If the metric was undefined, that is a modeling gap. If it existed and the model chose a neighbor, that is a description or selection problem. The fixes live in different places and different teams.
A rare question feeding a regulatory filing matters more than a common one feeding curiosity. Grade in layers and weight by what being wrong costs.9
Include questions that should not be answerable and assert that the system declines or clarifies. A suite that only contains answerable questions will happily pass a system that never refuses anything.
New metric, edited description, new model version. The suite is what lets you accept any of those in an afternoon instead of hoping.
A sequence that produces trust
The failure pattern is attempting to model the warehouse. It produces a year of work, no users, and a definition backlog nobody will adjudicate. The alternative is narrow and boring and it works.
Start from questions, not tables. Collect the twenty questions leaders actually ask, work out which metrics and dimensions they require, and model only those. Pick one domain with one clear owner who can settle definitional arguments in a day. Run read-only, with no write path and no actions, until the numbers are trusted. Publish the golden set and its pass rate where stakeholders can see it, because a visible score converts “do you trust the AI” into a conversation about specific failing cases. Then widen by evidence: the questions people asked that the layer could not answer are your backlog, in priority order, generated for free.
What you are actually building
The semantic layer is not middleware you add so that an LLM works. It is the written-down version of how your business measures itself, in a form a machine can execute and a human can review. Most organizations have never had that. The questions it forces — what is an active customer, which date is revenue measured against, who may see margin by region — are questions that were always unresolved, answered inconsistently in dozens of dashboards and spreadsheets.
That is why the work has value independent of any model. The layer makes BI consistent, makes definitional disagreements visible before they reach a board pack, and gives every consumer — a dashboard, a notebook, an agent — the same answer to the same question. The model is simply the consumer that makes the absence of such a layer impossible to ignore, because it will produce a confident number where a human analyst would have asked what you meant.
Do not ask the model to infer your business. Define it, publish it as a contract the model can only speak through, and keep the definitions in version control where they can be argued about by the people who own them. Then the interesting question stops being whether the model is smart enough and becomes whether your definitions are right — which is a question you can actually answer.
Sources and industry references
- Databright Cloud Solutions — Don’t Point the Model at the Warehouse: why raw data plus a capable model produces fluent answers rather than correct ones
- Cube — Semantic layer documentation: agents querying a warehouse directly produce inconsistent metrics and ungoverned access, and every query passes through the runtime where it is validated against the model and has access policies applied deterministically before reaching the warehouse
- Apache Ossie (incubating) — a declarative standard for defining metrics, dimensions, and joins so every tool and agent works from one source of truth, with a typed query surface that lets agents list, describe, and query models by name with self-correcting validation errors instead of writing SQL
- dbt Labs — MetricFlow: semantic models, primary and foreign entities as traversal paths, time and categorical dimensions, the simple, ratio, cumulative, derived and conversion metric types, identifier typing used to avoid fan-out and chasm joins, and definitions committed to version control for review
- Databright Cloud Solutions — The Temporal Truth Problem: separating when a change became effective from when a source recorded it and when a platform received it
- Databright Cloud Solutions — Redaction Is Not Governance: why authorization follows purpose rather than credentials, and what auditability requires
- Databright Cloud Solutions — The Model Is Not a Constant: retirement clocks, per-platform divergence, and why the layers around the model outlast it
- Snowflake — Open Semantic Interchange: the first specification version, January 2026, defining a vendor-neutral extensible model for data sets, metrics, dimensions, relationships and contexts under an Apache 2 license, intended to be interpreted consistently across tools, platforms and agentic applications
- Databright Cloud Solutions — Right Property, Right Time, Right Evidence: expected answers as specifications, layered grading, and weighting by consequence
This article provides architecture perspectives, not implementation guidance for a specific environment. Semantic layer tooling, specifications, and agent query interfaces are moving quickly; the specifics described here reflect published documentation as of September 2026, and Apache Ossie is an incubating project whose specification and interfaces may change. Validate the current modeling primitives, query APIs, and governance capabilities of your own platform before committing to a production design.