Semantic layer

ProductionRetrieval and dataPublished By Simon Budziak

A semantic layer is a software layer between company data and the tools that query it, defining business terms such as revenue or active customer once, so every dashboard, analyst and AI agent computes them the same way. It is how organizations stop AI from inventing its own definitions.

Why do AI systems need a semantic layer?

Because raw schemas do not say what words mean. Ask an AI agent for last quarter’s churn against bare tables and it will guess a definition, join plausibly and answer confidently, a numeric cousin of a hallucination. A semantic layer turns that guess into a lookup: the metric has one definition, the query compiles against it, and two agents asked the same question return the same number. Benchmarks of text-to-data systems keep finding the same thing: accuracy problems are mostly definition problems.

What does a semantic layer actually contain?

Definitions as code. The core objects are metrics (a measure plus its aggregation, grain and filters), dimensions to slice them by, entities with the joins that state how tables relate, and access policies deciding who may see what. A query names a metric and dimensions; the layer compiles it to SQL against the warehouse, so each definition lives once and executes everywhere. Because the definitions are files, they belong in version control, with review and tests on every change, which is what separates a semantic layer from a folder of agreed dashboard formulas that drift the day their author leaves.

How do AI agents query a semantic layer?

Through a structured API rather than raw SQL. The layer’s catalog of metrics and dimensions becomes the agent’s menu, exposed via tool calling or an MCP server: the model selects a metric, dimensions and filters as structured output, and the layer compiles and runs the query. The trade-off against free-form text-to-SQL is deliberate. Constraining the queries an agent can express removes the wrong-join failure mode, at the price of only answering questions the layer has definitions for, and an unanswerable request now fails loudly instead of plausibly.

Where does a semantic layer fit in an AI data stack?

Between the warehouse and everything that asks questions, next to retrieval. RAG grounds answers in documents; the semantic layer grounds them in metrics; a knowledge graph can hold the relationships both draw on. It is the most reusable piece of AI groundwork a data team can build, because every future agent inherits it. Building one is data work, not AI work, which is why it belongs inside data readiness for AI rather than inside any single project.

Frequently asked questions

What is the difference between a semantic layer and an ontology?

An ontology describes concepts and their relationships in general; a semantic layer is the operational version wired to your actual tables, with metric definitions, joins and access rules that queries execute against. Many semantic layers embed a small ontology; few ontologies can answer a query.

Is Snowflake or Databricks a semantic layer?

They are platforms a semantic layer sits on or ships with; both now offer semantic-model features, and standalone tools exist as well. The product matters less than the discipline: one agreed definition per metric, stored where every consumer, human or AI, must go through it.

Summarize this page with

Train your team to build this