Field Notes
Engineering

Dremio Reflections

How reflections speed up queries, and the types available.

1 min read

Reflection types ​

Reflections are physically optimized representations of source data. When they contain all the data needed by a query, they are used by Dremio to satisfy a query in place of source data.

  • If the data lake contains row-oriented data formats like CSV or JSON, a reflection will dramatically improve query performance over using the original raw data
  • Reflections store pre-computed results for later use. Complicated joins and data transformations can be computed in a reflection and reused, reducing the work that each query needs to perform
  • Reflections can partially or completely satisfy queries in place of source data, meaning reflections can be combined and reused in conjunction with other query filters, operations, and optimizations
  • Since reflections are an optional optimization transparent to users, you can add them iteratively when needed and tune your reflection strategy over time

Two key reflection types

  • raw reflection - One or more columns from the dataset, preserves row-level fidelity
  • Aggregate Reflection - Precomputed measures for all dimension combinations, like a materialized aggregate

image

  • Substitution refers to the ability of a reflection to satisfy a given query in place of raw physical data. When a reflection can substitute for raw data, we say it "covers" or all a portion of the query
  • The query planner compares the cost of using a reflection with the cost of executing the query against the physical data, and the lowest cost query plan is selected
  • Typically using one or more reflections is less expensive than executing the query against the raw physical data