Platform guide

Company signals as grounding data for Vertex AI

Which of Google's grounding options fits company signals, how to expose BigQuery tables to a Gemini model with typed queries, and how to return dates and sources so an answer can be checked.

Updated 5 October 20267 min read

Google now documents Vertex AI under the name Gemini Enterprise Agent Platform, and its name changes page lists the new names of its products: Vertex AI Search is now Agent Search, and Vertex AI Studio is now Agent Studio. This guide uses the current names, and the title keeps the name Vertex AI, which the module and host names in Google's own samples still carry: vertexai and aiplatform. It shows three ways to ground a Gemini model on Fokals company signals, a worked function-calling example over BigQuery tables, and the dates and sources to return so that an answer can be checked. Fokals is delivered direct, by REST API and as bulk files, which you load into BigQuery with its own load jobs, as the guide to loading company data into BigQuery describes.

What grounding means, and the options Google lists

Google defines grounding as connecting model output to verifiable sources of information. Its grounding overview says grounding reduces the chance that a model invents content, anchors responses to your data and provides links to sources. As of 1 October 2026 the overview lists these ways to ground a model: Google Search, Google Maps, Agent Search, RAG Engine, Elasticsearch, your own search API and several web-index options. None is named for BigQuery tables, so for company signals three routes matter.

RouteWhat Google's page saysWhere it fits company signals
Function callingThe model returns a structured call that names a tool and its parameters, your application runs the tool, and the result goes back to the modelCounts, scores and dated events held in tables
Grounding with your search APIGemini sends a query to an endpoint you provide, and the endpoint returns snippets, each with a uriA thin service in front of your BigQuery tables
Agent Search over a BigQuery tableA data store built from a table supports semantic search, and natural-language analytical queries are not supportedText fields such as announcement excerpts

Google also documents a remote BigQuery MCP server, whose execute_sql_readonly tool runs read-only SQL for a model. Its page lists a default limit of three minutes per query and 3,000 rows per result. The guide to serving company data to agents over Model Context Protocol covers that route.

Which signals are numbers and which are text

Most company signals are numbers with dates. A model that reads a row as a document cannot add, compare or window it, so let SQL do the arithmetic and let the model write the sentence. Text is the exception. The use case on retrieval over structured company signals weighs SQL against vectors in more depth.

DatasetWhat a row holdsRoute
Hiring Activity (company_hiring_daily), Sales Team Metrics (company_sales_weekly)Counts and medians per company and periodFunction calling
Intent Scores (company_intent_weekly)A score from 0 to 100 per topic, with the five strongest signals as evidenceFunction calling, with the evidence returned as sources
Technology Changes (company_tech_events)One dated change seen on a company websiteFunction calling
Company News (company_news)A title, a URL and at most 1,200 characters of the company's own textAgent Search, or the same function

The datasets are described in the data dictionary, and company data as context for AI agents covers tool design in general.

Function calling over BigQuery tables

Google's function calling page describes a two-step loop. You send the prompt with declarations of the functions the model may use, in the OpenAPI schema format. When the model needs a function it returns the name and the parameter values, your application runs the function, and you send the output back so that the model can write its answer. The declaration below is for one table. Every value reaches the query as a typed parameter, and Google's parameterized queries page says parameters cannot stand in for identifiers or table names, so the model can choose values and never the table.

{
  "name": "get_hiring_signals",
  "description": "Daily hiring counts for one company, up to a date, from Fokals tables. An empty list means no rows.",
  "parameters": {
    "type": "object",
    "properties": {
      "company_id": {"type": "string", "description": "The Fokals company_id."},
      "as_of": {"type": "string", "description": "Last day to include, YYYY-MM-DD."},
      "days": {"type": "integer", "description": "Closed days to return, 1 to 90."}
    },
    "required": ["company_id", "as_of"]
  }
}

The function body runs the query. The project name belongs to Acme Robotics, an invented company, so treat it as illustrative.

from datetime import date, timedelta
from google.cloud import bigquery

bq = bigquery.Client(project="acme-robotics-prod")

SQL = """
SELECT day, open_postings, new_postings, closed_postings, reconstructed
FROM `acme-robotics-prod.fokals.company_hiring_daily`
WHERE company_id = @company_id
  AND day BETWEEN DATE(@start) AND DATE(@as_of)
ORDER BY day
"""

def get_hiring_signals(company_id: str, as_of: str, days: int = 28) -> dict:
    days = min(max(days, 1), 90)
    start = date.fromisoformat(as_of) - timedelta(days=days - 1)
    config = bigquery.QueryJobConfig(query_parameters=[
        bigquery.ScalarQueryParameter("company_id", "STRING", company_id),
        bigquery.ScalarQueryParameter("start", "STRING", start.isoformat()),
        bigquery.ScalarQueryParameter("as_of", "STRING", as_of),
    ])
    rows = [
        {"day": r["day"].isoformat(), "open": r["open_postings"],
         "new": r["new_postings"], "closed": r["closed_postings"],
         "written_late": bool(r["reconstructed"])}
        for r in bq.query(SQL, job_config=config).result()
    ]
    return {"source": "Fokals company_hiring_daily", "rows": rows}

Four choices carry the weight. The company is a company ID (company_id), never a name: do the entity resolution in your application, so that the model cannot invent an ID. A small lookup table built from your loaded data is enough, with company_id, company, domain, ticker with MIC and ISIN. Match on domain or ISIN first and on name last, and ask the user to choose when two companies match. The result passes on reconstructed, which flags a period written more than seven days after it closed, as written_late, instead of hiding those rows. An empty list comes back as an empty list and never as zeros, so that absence is not read as no hiring. And the result names its source and the days it covers, so the answer can say what it is as of without guessing.

Grounding with your search API

Google's page on grounding with your search API describes a narrower contract. Gemini sends a POST request with a JSON body holding one query string to an endpoint you provide, authenticated with an API key sent as a query parameter named key. Your endpoint returns a JSON array of objects, each with a snippet and a uri, and an empty array when nothing matches. A request supports up to 10 grounding sources.

Because the request carries only free text, your service has to work out which company and which signal the query means, on every call. Function calling avoids that by declaring typed parameters. The search API route suits a team that already runs a search service in front of its data. An illustrative response for Acme Robotics follows.

[
  {
    "snippet": "Acme Robotics: 9 postings opened and 4 closed in the 7 closed days to 2026-10-03; 41 open on 2026-10-03. Source: company_hiring_daily.",
    "uri": "fokals:company_hiring_daily/acme-robotics/2026-10-03"
  }
]

Google describes uri as a link to the source or a relevant identifier, so an internal identifier is allowed. Where the row has a public address, such as the source link of a posting or an announcement, return that instead.

Retrieval over the text fields

Agent Search can build a data store from a BigQuery table. Google's data store page says it supports semantic search and does not support natural-language analytical queries, so it can answer which announcements mention a new plant and cannot answer how many postings opened last month. Periodic ingestion syncs every 1, 3 or 5 days, so an index can lag a daily table by that interval. The page also cautions that BigQuery permissions are not imported with the data, and that any user with enough Agent Search permissions can see the imported rows. Check that against the scope of your licence before you index licensed rows. The announcement excerpt is short, at most 1,200 characters of the company's own text, which suits retrieval. It sits in Company News, part of the announcements and scale dataset, which delivers structured rows and short excerpts, each with a link to the original. Full documents, such as contracts or complete filings, are the material the managed retrieval routes are built for.

What a good answer looks like

Take an illustrative question: how has hiring moved at Acme Robotics over the last four weeks? The application resolves the company to its company ID, the model calls get_hiring_signals with days set to 28, and the rows come back dated. A checkable answer reports the change in open_postings between the first and the last day, the balance of new_postings and closed_postings, and the last closed day it rests on. It says so when a row is marked as written late, and it describes what the counts show. Three instructions in the system prompt produce that behaviour.

  • State the dataset and the dates that every figure comes from.
  • Say that no rows were found when the function returns an empty list, and never write zero.
  • Describe what the data shows, with the dates it rests on.

Dates and sources

A grounded answer is only as checkable as the facts it carries. Return the date range, the dataset and, where one exists, the public source link of the posting, announcement or filing. For intent, return the evidence that comes with each Intent Scores row: up to five signals, each with its date, kind, source and weight, so the model can name what drove a score. Every score carries the dated signals behind it, as the methodology sets out, and the model's instructions should say so. The intent dataset describes the score.

Fokals data is company-level, and a leadership change is recorded by role. Datasets refresh daily and weekly, so an answer is as current as the last closed day, and the answer should state that day. Google's overview says grounding reduces the chance of invented content, and reducing is not removing, so keep a check on answers that matter.

Frequently asked questions

What is Vertex AI called now?

Google's documentation now describes Gemini Enterprise Agent Platform, and its name changes page lists the new product names. Vertex AI Search is Agent Search, Vertex AI Studio is Agent Studio, and Vertex AI Agent Engine is Agent Runtime. Module and host names in Google's code samples still use vertexai and aiplatform.

Can Gemini be grounded on BigQuery tables?

Yes, by three documented routes. Function calling lets your application run a parameterised query when the model asks for one. Grounding with your search API lets Gemini call an endpoint you provide, which can query BigQuery. Agent Search can build a data store from a BigQuery table, though Google says it supports semantic search and not natural-language analytical queries.

Should company signals go into a vector index?

Only the text. Announcement excerpts and titles suit retrieval by meaning. Counts, scores, rates and dates need arithmetic and filters by company and date, which SQL does exactly and an embedding does not. Keep the numbers in BigQuery behind a function, and index the text if you index anything.

How do I make a grounded answer cite its sources?

Return the source with every fact: the dataset, the dates it covers and, where the row has one, its public source link. Google's overview says grounding provides auditability through links to sources, and a search API result carries a uri for that purpose. For an intent score, return the evidence so that the answer can name the signals behind it.

How do I ground a Gemini model on Fokals data?

Fokals is delivered direct, by REST API and as bulk files. Load the datasets into BigQuery with its own load jobs, then expose them to the model through one of the routes above, most often a function that runs a typed query and returns dated rows with their sources. The data dictionary lists every dataset and field you can return.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.

What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.