Platform guide

Using company signals with Snowflake Cortex AI functions

Which Cortex functions exist today, worked SQL to summarise a watchlist's announcements and to search them, and the governance, cost and licence questions to settle first.

Updated 5 October 20266 min read

This guide uses Snowflake Cortex to summarise and search the announcements in Company News for a watchlist of companies, with every sentence traceable to a dated source. It also covers who may call the functions, where requests are processed, what they cost and what a licence has to say about the output. Fokals is delivered direct, by REST API and as bulk files, which you load into Snowflake with its own loaders, as the guide to loading company data into Snowflake shows. The announcements are part of the announcements and scale dataset.

What Cortex offers, as documented on 4 October 2026

Snowflake's page on Cortex AI Functions lists these SQL functions: AI_COMPLETE, AI_CLASSIFY, AI_FILTER, AI_AGG, AI_EMBED, AI_EXTRACT, AI_SENTIMENT, AI_SUMMARIZE, AI_SUMMARIZE_AGG, AI_SIMILARITY and AI_TRANSLATE, among others. The older SNOWFLAKE.CORTEX.* names such as COMPLETE and SUMMARIZE are legacy. Snowflake's migration page maps each to an AI_* function and calls the AI_* functions the primary interface and the focus of future development.

NeedCortex featureFokals datasetWhat to know
Summarise a company's announcementsAI_AGG, AI_SUMMARIZE_AGGCompany News (company_news)Accept more text than a model's context window
Summarise one textAI_SUMMARIZECompany NewsOne input, one summary
Find announcements by meaningCortex SearchCompany NewsA service over a table, filtered by attributes
Ask a question about numbersCortex Analyst, or Cortex AgentsHiring Activity (company_hiring_daily) and othersNeeds a semantic view that you define
Label an announcementAI_CLASSIFYnone neededThe event types are already a labelled field

Use a model only for text you need to read. Fokals already labels each announcement with zero or more of 13 event types, so filtering on the event type is a query and not a model call. Questions about numbers, such as which companies opened the most sales roles, are exact in SQL on Sales Team Metrics (company_sales_weekly). Cortex Analyst generates SQL from a question over a semantic view, and its page now points to Cortex Agents, which list Cortex Analyst and Cortex Search among their tools. That is a layer to add after the tables are sound, and the use case on retrieval over structured company signals discusses when SQL beats search.

Summarise a company's announcements

AI_AGG takes a text expression and an instruction and, according to its reference page, supports datasets larger than a model's context window, so a month of announcements per company is one aggregate query. AI_SUMMARIZE_AGG takes only the text and returns a general summary; its page says to use AI_AGG for a more specific one. The query below builds a digest for each company on a watchlist table.

select n.company_id,
       any_value(n.company) as company,
       count(*)             as announcements,
       ai_agg(
         concat_ws(' | ', to_char(n.at, 'YYYY-MM-DD'), n.title, coalesce(n.excerpt, ''), n.url),
         'Each input is one announcement published by the same company: date, title, excerpt and address, separated by bars. Write at most five lines. Each line states one fact that the text gives and ends with the date and address of its announcement. Do not add anything the text does not state.'
       ) as digest
from fokals.company_news n
where n.company_id in (select company_id from watchlist)
  and n.at >= dateadd('day', -30, current_date())
group by n.company_id;

Three choices in that query carry the weight. The input carries the date and the address of each announcement, so the instruction can require them and a reader can open the source. concat_ws is wrapped in coalesce because Snowflake's CONCAT_WS returns NULL when any argument is NULL, and an announcement with no excerpt would otherwise vanish. The instruction forbids additions, which lowers the risk of invented detail and does not remove it, so sample the digests against the sources before anyone relies on them.

The excerpt field holds at most 1,200 characters of the company's own text. For leadership changes Fokals records the role concerned, so ask the model to describe a change by role too. AI_COMPLETE is the function to use when you need output in a fixed shape: it takes a model name from the list for your region and a response format as a JSON schema. Model lists change, so read the page on models and regional availability before you fix a name in a pipeline.

The same method explains an intent surge. A surge in Intent Scores (company_intent_weekly) from the intent dataset carries its five strongest signals as evidence, each with a date, kind, source and weight. Pass those rows to AI_AGG with an instruction to list them in order of weight and add nothing, and a salesperson gets the reasons behind the number in plain words. The model rewords the evidence and never scores it: the score and the surge flag are computed under the frozen version intent-v2, and nothing a model writes should feed back into them.

Grade the output as Fokals grades its own labels. Show a person the exact input and mark each line of the digest correct, wrong or unsure, then divide correct by correct plus wrong. Do it on a sample before the digest reaches anyone, and again whenever you change the instruction or the model.

Search the announcements

Cortex Search is described by Snowflake as hybrid search that combines vector and keyword matching with semantic reranking over a table. A service needs a source query, one search column, optional attribute columns that can be returned and used as filters, a warehouse and a target lag. Change tracking must be on for the underlying objects. Snowflake recommends search text of no more than 512 tokens, about 385 English words, and an excerpt of at most 1,200 characters is roughly 200 words.

alter table fokals.company_news set change_tracking = true;

create or replace cortex search service fokals.announcement_search
  on search_text
  attributes company_id, company, ticker, url, title, event_types, published_at
  warehouse = search_wh
  target_lag = '1 day'
as (
  select n.company_id, n.company, n.ticker, n.url, n.title,
         n.at::date           as published_at,
         n.event_types::array as event_types,
         concat_ws('. ', n.title, coalesce(n.excerpt, '')) as search_text
  from fokals.company_news n
);

select parse_json(
  snowflake.cortex.search_preview(
    'fokals.announcement_search',
    '{
       "query": "chief financial officer appointed",
       "columns": ["company", "published_at", "title", "url"],
       "filter": {"@and": [
         {"@contains": {"event_types": "leadership_change"}},
         {"@gte": {"published_at": "2026-09-01"}}
       ]},
       "limit": 10
     }'
  )
)['results'] as results;

The filter uses documented operators: @contains for an array column and @gte for a date. Snowflake says SEARCH_PREVIEW is for testing and validation and not for serving an application, which should use the Python or REST API. Creating a service needs the CREATE CORTEX SEARCH SERVICE privilege on the schema, SELECT on the source and one of the SNOWFLAKE.CORTEX_USER or SNOWFLAKE.CORTEX_EMBED_USER database roles, and querying needs USAGE on the service, database and schema. The cost page for a service names warehouse compute for refreshes, embedding tokens, serving compute, storage and cloud services compute.

Govern access, processing and cost

By default everyone can call the AI functions. The privileges page says a caller needs the USE AI FUNCTIONS account privilege and one of the CORTEX_USER or AI_FUNCTIONS_USER database roles, and that the privilege and CORTEX_USER are both granted to PUBLIC by default. AI_FUNCTIONS_USER is not, and it excludes the aggregate functions AI_AGG and AI_SUMMARIZE_AGG, so the digest above needs CORTEX_USER. Older guides restrict models with CORTEX_MODELS_ALLOWLIST. The same page says it is being deprecated and that from August 2026 it can only be set to 'None', and it describes controlling model access by granting application roles for model objects instead.

-- licensed data should not be open to every user of the account
revoke database role snowflake.cortex_user from role public;
grant  database role snowflake.cortex_user to role analyst_ai;

-- where requests may be processed
show parameters like 'CORTEX_ENABLED_CROSS_REGION' in account;

-- what the functions cost, by function and model
select date_trunc('day', start_time) as day, function_name, model_name, sum(credits) as credits
from snowflake.account_usage.cortex_ai_functions_usage_history
where start_time >= dateadd('day', -30, current_timestamp())
group by 1, 2, 3
order by 1 desc, 4 desc;

Processing location is a setting. Snowflake's cross-region inference page says the CORTEX_ENABLED_CROSS_REGION parameter, which only ACCOUNTADMIN can change, decides whether requests may leave the account's home region. DISABLED keeps them in the home region, and the page says customer data stays stored in your region while the prompt and response are sent transiently to the processing region. Snowflake's AI overview states that its AI models run inside its security and governance perimeter and that it does not use customer data to train models made available across its customer base.

Cost follows tokens. The cost page says text generation functions bill both input and output tokens and recommends a warehouse no larger than MEDIUM. Test on a handful of companies first, count tokens with AI_COUNT_TOKENS, and watch the usage view above.

What you may do with the output depends on the licence. Decide which of internal use, embedding in a product and redistribution a feature falls under before you build it, and have the agreement say so. A digest read by your own analysts is one use. A digest shown to your customers is another, and Snowflake lists Cortex Search services among the objects a share can carry, so a service shared to other accounts raises the question of redistribution. The guide to Secure Data Sharing for licensing teams sets out those terms.

Where this stops

A model summarises the text it is given, so the digest is as sound as the rows behind it. Company News carries company announcements and regulatory disclosures, classified into 13 event types under named, frozen versions, and each item links to its original. Use a filter on the event types as a screen over those classifications, and read the source for a verdict.

For hiring, Job Postings delivers titles, labels and the tools named, which are better counted in SQL than summarised. Keep each digest with its input date range, the label version of the data and the function and model that made it, because the same query tomorrow will cover other days and may read differently. For a back-test, summarise only text that existed on the date being tested. The methodology states how labels are produced and versioned.

Frequently asked questions

What are Snowflake Cortex AI functions?

They are SQL functions that call language models from inside Snowflake, among them AI_COMPLETE, AI_SUMMARIZE, AI_AGG, AI_CLASSIFY, AI_EXTRACT and AI_EMBED. Snowflake's documentation says the models are deployed within the Snowflake Service perimeter. The older SNOWFLAKE.CORTEX names, such as COMPLETE and SUMMARIZE, are legacy and map to AI_ functions.

Can I use Snowflake Cortex on data I license from a vendor?

Technically yes, once the data is in your account. Whether you may is a licence question. Summaries read by your analysts are one use, features shown to your customers are another, and a search service shared to other accounts is a third. Decide which of internal use, embedding in a product or redistribution applies, and have the agreement name it.

How do I search company announcements with Cortex Search?

Turn on change tracking for the table, then create a Cortex Search service with the announcement text as the search column and the company, date and event types as attributes. Query it with filters such as the array filter for an event type. SEARCH_PREVIEW is for testing, and an application should use the Python or REST API.

Does Snowflake train models on my data?

Snowflake's documentation says it does not use customer data to train models that are made available across its customer base, and that its AI models run inside its security and governance perimeter. Your agreement with Snowflake is the binding text, so read it with the documentation page open.

Who can use Cortex AI functions in my Snowflake account?

By default everyone: the USE AI FUNCTIONS privilege and the CORTEX_USER database role are both granted to PUBLIC. To restrict use, revoke them from PUBLIC and grant them to named roles. AI_FUNCTIONS_USER is not granted to PUBLIC and covers the scalar functions but not AI_AGG or AI_SUMMARIZE_AGG.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.

What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.