How to Cut Prompt Costs by 90% with Prompt Caching in Snowflake

Uli Bethke

Uli has been rocking the data world since 2001. As the Co-founder of Sonra, the data liberation company, he’s on a mission to set data free. Uli doesn’t just talk the talk—he writes the books, leads the communities, and takes the stage as a conference speaker.

Any questions or comments for Uli? Connect with him on LinkedIn.


Published on October 2, 2026

When building production LLM applications—whether question-answering systems over voluminous documentation, conversational agents with sprawling context windows, or batch processing pipelines—input token processing dominates both latency and compute costs.

To solve this, Snowflake supports Prompt Caching. On paper, it promises up to a 90% discount on cached input tokens. But where can you actually use it? How does it behave under the hood? How do you implement it across different model families, and how can you reliably monitor your cache hits when the API responses themselves can be deceiving?

In this post, we’ll explore how prompt caching works in Snowflake, unpack core concepts like ephemeral caching and cache breakpoints, clarify supported and unsupported Snowflake services, walk through a reproducible test using the claude-sonnet-4-5 model, and review the gotchas you must know before going to production.

TL;DR 

  • The Promise: Prompt caching in Snowflake Cortex reduces input token costs by 90% and significantly improves latency for repetitive prompts.
  • Availability: Prompt caching is available in Cortex REST API, AI Gateway. It is not available in AI_COMPLETE and other AI functions
  • The Rules: To qualify, the cached block must be at least 1,024 tokens. Caching requires an exact byte-for-byte prefix match from the very first token; any dynamic data placed at the top of the prompt will invalidate the cache.
  • The Gotchas: The direct Cortex REST API response does not break out the cache reads and writes and does not include billing information. Built-in SQL functions (like AI_COMPLETE) do not support caching at all.
  • The Solution: Route requests through the Cortex AI Gateway to gain full, per-request telemetry including cache reads / writes and per request billing info.

Overview: What is Prompt Caching and Why Care?

Every request sent to a large language model requires the transformer engine to ingest, tokenize, and compute attention keys and values (the KV cache) across all prompt tokens before emitting a single output token. In enterprise workflows, applications routinely resend identical blocks of text:

  • Long system prompts and enterprise guardrails
  • Extensive API tool descriptions
  • Large enterprise reference manuals or policy documents
  • Accumulated conversational turns

Without caching, you pay in both compute credits and latency to re-process that exact same prefix on every single call.

The Value Proposition

  • Cost Reduction (90% Off): In Snowflake Cortex, cached input tokens that result in a cache hit (read) are billed at only 10% of the regular input rate (a 90% discount), provided the cached block is at least 1,024 tokens. Cache writes are billed at the standard input rate.
  • Latency Improvements: Reusing pre-computed KV representations skips initial prompt computation. While latency savings for small prompts (~2,000 tokens) can be masked by normal network jitter, the throughput and time-to-first-token (TTFT) gains grow significantly on larger context blocks (10k+ to 100k+ tokens).
  • Ideal Fits: RAG pipelines querying the same static source text, chatbots maintaining long message histories, and multi-tenant batch jobs sharing static enterprise context.
  • Poor Fits: Single one-off queries, prompts smaller than 1,024 tokens, requests spaced farther than 5 minutes apart, or prompts where the beginning of the text changes dynamically on every call.

Support in Snowflake: Where Prompt Caching Works (and Where It Doesn’t)

A common point of confusion is assuming prompt caching applies globally across all Snowflake generative AI features. It does not.

Supported Surfaces

  • Cortex REST API: Full support across both the /chat/completions and /messages endpoints.
  • Cortex AI Gateway (Cortex Gateway AI): Supported when routing LLM requests through Snowflake’s centralized gateway infrastructure, enabling unified policy control and caching benefits for external clients.

Unsupported Surfaces

  • SQL AI Functions (e.g., AI_COMPLETE, SNOWFLAKE.CORTEX.COMPLETE): Built-in SQL functions do not support prompt caching. When invoking AI_COMPLETE() in SQL queries or stored procedures, prompts are evaluated statelessly without prefix caching or cache breakpoint markers. If your architecture relies on caching large shared contexts, you must route those requests through the Cortex REST API or Cortex AI Gateway rather than invoking SQL functions.

Core Concepts: “Ephemeral” and “Cache Breakpoints”

To implement prompt caching effectively, two technical concepts must be understood: the ephemeral lifecycle and cache breakpoints.

What Does “Ephemeral” Mean?

In the context of LLM prompt caching (originating from Anthropic’s caching architecture and adopted by Cortex), ephemeral refers to the cache’s temporary, volatile lifecycle:

  • Volatile In-Memory Storage: The cached KV states live in high-speed, volatile GPU/accelerator memory rather than persistent storage.
  • 5-Minute Time-To-Live (TTL): A cache entry remains alive for exactly 5 minutes after creation.
  • Sliding Window: Every subsequent request that successfully performs a cache read automatically resets the 5-minute countdown timer.
  • Automatic Purging: If 5 minutes elapse without any requests referencing that exact prefix, the cache entry is purged automatically. Currently, “ephemeral” is the only cache type supported by Cortex.

What are Cache Breakpoints?

Large language models process text strictly sequentially from left to right. Therefore, prompt caching is prefix-based: any cache hit requires an exact, byte-for-byte match starting from token index zero up to a designated cut-off point.

A cache breakpoint is an explicit marker telling the caching engine: “Process everything up to this exact point, store the resulting state in the cache, and evaluate future requests against this prefix.”

Key rules around breakpoints:

  1. Claude Models: Require explicit breakpoint definition via the cache_control: {“type”: “ephemeral”} parameter within the message content structure.
  2. OpenAI Models: Cache implicitly without explicit breakpoints on matching prefixes exceeding 1,024 tokens.
  3. Breakpoint Limits: You can define a maximum of 4 cache breakpoints per request. This allows multi-tiered caching hierarchies (e.g., common base prompt domain context user session).

Real-World Scenarios: Where Prompt Caching is Most Beneficial

Prompt caching is highly effective for workloads where a large, static block of text serves as the foundation for multiple subsequent queries. Common architectures that benefit include:

  • Conversational Agents: Maintaining context windows with extensive system instructions, tools, and message history.
  • Batch Processing: Multi-tenant jobs or data engineering pipelines sharing a static enterprise context.

A Practical Example: The 700-Column Data Mapping Pipeline

To illustrate the architectural impact, consider a recent data integration project. The objective was to map 150 standardized target columns to a legacy schema containing 700 source columns. To ensure accurate mapping and minimize hallucinations, the LLM required the full semantic context for all 700 source columns, which included business definitions, expected data types, and sample values.

Compiling this metadata resulted in a large token payload.

The Uncached Approach Processing this traditionally would require sending the full 700-column context repeatedly to evaluate the target columns. This approach incurs high input token costs and latency bottlenecks, as the transformer must re-evaluate the identical schema for every call.

The Caching Solution Using Cortex prompt caching, we set the schema text as our static prefix and applied a cache breakpoint (cache_control: {“type”: “ephemeral”}). We paid the standard input token rate once to load the schema into memory.

Rather than sending 150 individual requests, we batched the target columns into groups of five. For each request, we appended the small, variable batch to the end of the prompt:

“Based on the cached schema above, which source columns best map to the following 5 target columns: [customer_id, order_date, total_amount, status, lifetime_value]?”

The Result The initial call triggered the cache write, and all subsequent batched calls registered as cache reads. This provided a 90% discount on input token costs for the remainder of the process and significantly reduced the latency of each response.

How to Use Prompt Caching in Snowflake Cortex REST API

Rules of Engagement

  1. Minimum Block Size (1,024 Tokens): The cached prefix must be at least 1,024 tokens (~4,000 characters). Anything shorter bypasses the cache entirely with no warning.
  2. Exact Byte-for-Byte Prefix Matching: The cached block must be identical down to every whitespace, newline, and punctuation mark starting strictly from the first byte of the prompt.
  3. Model Selection: Ensure you are using modern models that support caching (e.g., claude-sonnet-4-5, claude-3-7-sonnet, gpt-4o). Deprecated versions like claude-3-5-sonnet will return HTTP 400 errors on Cortex REST.

Test Case Design & Reproducible Python Script

To verify caching behavior, billing accuracy, and expiration mechanics, we designed a reproducible test scenario:

  1. Call 1 (Cache Write / Miss): Send a large static block (~2,023 tokens, well above the 1,024-token floor) marked with cache_control: {“type”: “ephemeral”}, followed by Question A.
  2. Call 2 (Cache Read / Hit): Send the exact same static block followed by Question B.
  3. Call 3 (Cache Expiry Verification): Wait 5+ minutes, re-send the identical payload, and verify whether a fresh write occurs.

Below is the complete, production-grade test script using Snowflake Key-Pair authentication.

Test Outcomes & Monitoring Verification

Why You Can’t Trust the API Response Body

If you inspect the usage object returned by the Cortex API response, you will encounter confusing metrics:

Notice that cache_write_tokens is 0, while cached_tokens reports 2023 on both Call 1 and Call 2. The API response does not split reads from writes—it combines both under cached_tokens.

The True Source of Truth: ACCOUNT_USAGE

To verify what actually happened at the billing layer, query SNOWFLAKE.ACCOUNT_USAGE.CORTEX_REST_API_USAGE_HISTORY using the IDs captured from the x-snowflake-request-id response header:

Verified Test Results

Run PhaseCallRequest IDUsage View Granular MetricsAPI Response ClaimReality
Initial RunCall 18921aa64…cache_write_input: 2023, input: 10cached: 2023, write: 0Cache Write (Full Price)
Initial RunCall 244e3095a…cache_read_input: 2023, input: 11cached: 2023, write: 0Cache Read (90% Discount)
After 5+ MinCall 14cca220a…cache_write_input: 2023, input: 10cached: 2023, write: 0Cache Write (Expired TTL)

The account usage view confirms the mechanism:

  • Call 1 created a 2,023-token cache entry (cache_write_input).
  • Call 2 reused those tokens at the discounted tier (cache_read_input).
  • After waiting 5 minutes without activity, the cache dropped, and the identical request triggered a fresh write.
Database log table displaying request IDs, total token counts, and JSON breakdown of granular cache read and write tokens.

Prompt Caching in Cortex AI Gateway

Relying on hidden headers and asynchronous billing views to track prompt caching efficiency is a cumbersome workaround for production monitoring. Fortunately, the recent release of the Cortex AI Gateway natively resolves the standalone REST API’s telemetry limitations. By acting as a central control plane for enterprise LLM requests, the Gateway inherently supports and tracks prompt caching workloads with full transparency.

  • Centralized Proxy: It unifies authentication, rate limiting, guardrails, and observability across external tools, client applications, and internal Cortex models.
  • End-to-End Telemetry: It captures full request lifecycle events and routes detailed metrics directly into account telemetry views.
  • Seamless Prompt Caching Observability: Because all requests and telemetry now route through this single gateway, it supports and tracks the prompt caching workloads that the standalone REST API could not manage.

Enhanced Observability & Snowsight Integration

The Cortex AI Gateway records OpenTelemetry traces for every inference request. These traces provide full execution transparency: which models were invoked, token consumption and status. 

LLM inference trace screen displaying cache read hit metrics, token counts, and execution details for Claude Sonnet.
LLM inference trace dashboard displaying request metrics, token counts, and execution details for Claude Sonnet.

Each gateway request produces one span. The Observability UI (AI & ML >Cortex AI Gateway > Observability) shows a trace detail panel with the following fields:

Trace identity:

  • Trace ID — unique identifier for the end-to-end execution. All spans in the same agent turn share this ID.
  • Span ID — identifier for this specific inference call within the trace.
  • Request ID — HTTP request identifier. This is the join key to AI_GATEWAY_USAGE_HISTORY for cross-referencing traces with credit consumption.

LLM Inference metrics:

  • Model name — the model that served the request (e.g., claude-sonnet-4-5).
  • Token count — total tokens processed (input + output).
  • Input tokens — tokens sent to the model. Includes cached tokens —they count as input but are read from cache rather than reprocessed.
  • Output tokens — tokens generated by the model.
  • Cache write tokens — tokens written to the prompt cache on this call.Present on the first call that includes a cache_control block.
  • Cache read tokens — tokens served from an existing cache. Present on subsequent calls with the same cached prefix.
  • Status — Success or Error.

Step by Step process of testing AI Gateway Cache Behaviour

Run the following steps directly in the terminal 

Step 1: Set your variables

Step 2: Build the cached context

This generates the large text block and writes both request bodies to files:

Step 3: Send call 1 (expect cache write)

Step 4: Send call 2 immediately (expect cache read)

Step 5: Query the usage view

Data table showing AI inference token usage, cache read and write stats, and credit costs for Claude Sonnet.

Example output of AI Gateway Prompt Caching Test

Unlike the direct REST API endpoint—which folds write tokens into cached counts and always returns 0 for cache_write_tokens—the AI Gateway returns fully transparent token usage objects in its HTTP response body:.

Empirical Test Results & Telemetry Breakdown

Executing the same reproducible test scenario using claude-sonnet-4-5 with a 2,025-token cached block through the Cortex AI Gateway gave the following telemetry in SNOWFLAKE.ACCOUNT_USAGE.AI_GATEWAY_USAGE_HISTORY:

  • Call 1 (Cache Write): Recorded cache_write_input_tokens: 2025 with a total credit cost of 0.003962 credits.
  • Call 2 (Cache Read): Recorded cache_read_input_tokens: 2025 with a total credit cost of 0.000469 credits, an 88% reduction in credit costs compared to the initial write call.
  • Cost Savings: The combined cost for this request pair was 0.004431 credits, compared to ~0.008 credits for two uncached executions. That is an immediate 44% cost reduction on just a two-prompt sequence. At scale, these savings approach 90% for large batch workloads.

Key Observability Improvements vs. Cortex REST API

  • Explicit Token Categorization: Telemetry views provide distinct, dedicated fields for cache_write_input_tokens and cache_read_input_tokens, eliminating the confusion of combined metrics.
  • Direct Per-Request Credit Costing: Exact credit burn is directly attributed to each individual request ID in the credits_consumed column of AI_GATEWAY_USAGE_HISTORY. In contrast, the direct Cortex REST API does not expose per-request credits in CORTEX_REST_API_USAGE_HISTORY.

Critical Prompt Caching Gotchas to Keep in Mind

Byte-for-Byte Prefix Rule: Dynamic Data at the Top Destroys the Cache

LLM prompt caching operates strictly from token 0 sequentially forward. Any alteration in the prefix—a dynamic timestamp, a session UUID, user metadata, or even minor whitespace changes at the top of the prompt—invalidates all cached blocks downstream.

Anti-pattern (Cache Bust Every Time):

Best Practice:

Always structure your prompts with invariant, reusable content first, and append variable data (timestamps, queries, user context) after the cache breakpoint.

No Prompt Caching in SQL Functions (AI_COMPLETE)

Prompt caching is not supported in built-in SQL functions like AI_COMPLETE or SNOWFLAKE.CORTEX.COMPLETE. If you execute SQL queries containing large static contexts, every query is treated as an uncached, full-price execution. You must use the Cortex REST API or Cortex AI Gateway to get caching.

No Per-Call Credit Visibility in Cortex REST API

While CORTEX_REST_API_USAGE_HISTORY gives granular token counts (cache_write_input, cache_read_input), Snowflake does not expose per-request credit consumption for the direct REST API. Account-level credits aggregate in general billing views, so you must manually multiply granular token counts by Snowflake model credit rates. However, routing requests through the Cortex AI Gateway resolves this limitation by populating the credits_consumed column directly in AI_GATEWAY_USAGE_HISTORY.

Request IDs Are in the Headers, Not the Body

The JSON response body returned on successful 200 OK calls leaves the id field empty. The actual identifier needed to join against ACCOUNT_USAGE is returned exclusively in the x-snowflake-request-id HTTP header. If your client does not record response headers, you will have no reliable way to trace individual calls back to Snowflake billing records.

API Response Mislabels Writes as Reads

As shown above, cache_write_tokens is unsupported in the REST response and consistently outputs 0. The API folds writes into cached_tokens. Do not build internal metering or telemetry solely on the HTTP response payload; rely on the Account Usage views.

Silent Failures Below Threshold

If you submit a prompt that is 1,023 tokens long with cache_control: {“type”: “ephemeral”}, Cortex does not raise an exception, nor does it return a warning. It silently executes the model call, bills you the full standard input token rate, and ignores the cache directive.

No Manual Cache Invalidation Endpoint

Snowflake provides no API command or SQL function to purge or invalidate a prompt cache on demand. To force a cache miss during testing or document updates, you must either wait out the 5-minute TTL or prepend/modify a dynamic token (such as a UUID or version tag) at the start of your prompt.

Latency Variance on Small Payloads

Do not rely on wall-clock latency to confirm whether caching is functioning. At ~2,000 tokens, prompt evaluation savings (~50–100ms) are easily obscured by network latency and model generation variance. Rely strictly on token breakdown metrics.

Uli Bethke

About the author:

Uli Bethke

Co-founder of Sonra

Uli has been rocking the data world since 2001. As the Co-founder of Sonra, the data liberation company, he’s on a mission to set data free. Uli doesn’t just talk the talk—he writes the books, leads the communities, and takes the stage as a conference speaker.

Any questions or comments for Uli? Connect with him on LinkedIn.