Dremio Reflections
How reflections speed up queries, and the types available.
Reflection types
Reflections are physically optimized representations of source data. When they contain all the data needed by a query, they are used by Dremio to satisfy a query in place of source data.
- If the data lake contains row-oriented data formats like CSV or JSON, a reflection will dramatically improve query performance over using the original raw data
- Reflections store pre-computed results for later use. Complicated joins and data transformations can be computed in a reflection and reused, reducing the work that each query needs to perform
- Reflections can partially or completely satisfy queries in place of source data, meaning reflections can be combined and reused in conjunction with other query filters, operations, and optimizations
- Since reflections are an optional optimization transparent to users, you can add them iteratively when needed and tune your reflection strategy over time
Two key reflection types
- raw reflection - One or more columns from the dataset, preserves row-level fidelity
- Aggregate Reflection - Precomputed measures for all dimension combinations, like a materialized aggregate
- Substitution refers to the ability of a reflection to satisfy a given query in place of raw physical data. When a reflection can substitute for raw data, we say it "covers" or all a portion of the query
- The query planner compares the cost of using a reflection with the cost of executing the query against the physical data, and the lowest cost query plan is selected
- Typically using one or more reflections is less expensive than executing the query against the raw physical data