In a day in the life of a SQL query, we followed a single query from question to answer. Now we're looking at what happens when that query, runs across Chats, Live Apps, and Dashboards. Here's how WisdomAI uses caching to deliver fast answers in chat and Live Apps without sacrificing data freshness.
Why we need caching
Consider what happens when someone opens a Live App. They'll likely change a filter, then drill into a metric, and each of those steps sends a new query to the warehouse. Now multiply that by every user who opens the same app. The queries are going to pile up, users are left waiting, and a single app can fan out into a a large number of SQL queries hitting the warehouse, all at once.
WisdomAI connects to a range of SQL databases and analytical warehouses, and each comes with its own limitations. Some need time to wake up after sitting idle, others cap how many queries can run at once, and some simply aren't sized for heavy traffic. On top of that, every repeated query adds to both the workload and the token bill.
This is where caching comes in. By reusing data that has already been fetched, we can return answers faster while sending fewer queries to the warehouse, which also brings compute costs down.
Of course, warehouses and databases have their own caches, but those typically rely on syntactic matching. Because WisdomAI has full visibility into how queries are used and what they mean, we can go further with semantic caching, which increases cache efficiency severalfold.
How Wisdom caches queries
WisdomAI uses 2 caching levels: one for identical queries, and one for queries that overlap.
Query result caching
A query that starts in a chat often reappears later in a dashboard or starter question. Wisdom stores every query result as a Parquet file. When a structurally identical query runs within the configured cache interval, it reuses that stored result instead of sending the query back to the warehouse.
Data freshness is the obvious concern here, so we also run table modification checks to see whether the source data has changed since it was last fetched. If it hasn't, the stored result can still serve an up-to-date answer.
View materialization
Result caching works well for exact matches, but many queries are similar rather than identical. To cover these, we allow tables and derived views to be materialized inside WisdomAI on a regular cadence.
When a query comes in, our matching algorithm determines whether it can use a single cached view or a combination of cached views. If there’s a match, WisdomAI executes the query against the stored data and serves the result directly, without going back to the warehouse.
Making this work depends on how those cached views are built and read, which is what the next section covers.
Cached views
Building cached views
Configuring cached views can be done in two ways:
Users select domain tables or derived tables and configure refresh cadence themselves.
Users choose to optimize a slow-loading app, and WisdomAI builds cached views automatically in the background.
Once we have the configured SQL queries for each connection, our internal job framework divides the work into smaller activities. These activities stream data from the warehouse into our cached storage.
For that storage, we chose Iceberg as the open table format, with Parquet as the underlying file format, and keep the data in cloud object storage for scalability and durability.
Because warehouse datasets can be large, we also apply safeguards while building the cache. We use relevant projections and filters to limit how much data comes over, and discard any candidates that fall outside the supported size limits.
Reading cached views
With the cached data is in place, the next step is deciding which cached view, or combination of views, can serve an incoming query.
When all the base tables used by a query are cached, we prefer to read from those cached copies by replacing the source references with our internal storage paths. For example, a query against a source table:
can run against its cached data instead:
We can also go a step further by rebasing a query onto a suitable cached view. This lets us reuse data that has already been filtered or aggregated, then run only the remaining operations inside WisdomAI.
To see how that works, consider a cached view that holds total sales by region and product category:
Now an incoming query asks for total sales by region, but only for the Electronics category:
Since the cached view already contains everything this query needs, WisdomAI can answer it by applying the category filter and summing the stored totals:
The filtering and aggregation also happen inside WisdomAI, without querying the source database again.
Executing queries inside Wisdom
Wisdom SQL Dataframe goes through a series of transformations before execution, as explained here. One of these is the cache rewrite transformation. When it finds a suitable cache match, it rewrites the SQL to read from the cached data.
That rewritten query then runs inside WisdomAI on an embedded SQL engine. The engine reads the data it needs from object storage and keeps those byte ranges on local disk, so repeat reads are faster. To keep performance predictable, each pod runs a controlled number of concurrent queries, and we add more pods as the workload grows.
Caching is not easy
All of this sounds straightforward on paper, but a few challenges can make it harder in practice.
The first is that WisdomAI accepts LLM-generated SQL, so queries can vary widely in structure and dialect. To run them on our engine, we need to translate (or “transpile”) them into Wisdom’s SQL dialect while keeping them semantically equivalent.
The second is that when we build a cached view, we can't know in advance which filters or drilldowns users will apply later. The view needs to retain enough data to support future filters and drilldowns. If a later query needs data the view does not contain, it must go back to the source database.
For users, none of this complexity is visible. A Live App simply loads faster, and the warehouse sees far fewer repeated queries. If you're building something similar, we'd love to hear how you approach it.



