Use case

Company data as context for AI agents and copilots

Design the tools, the result format and the tests for an agent that answers questions about companies from dated, sourced rows instead of scraped pages.

Updated 5 October 20267 min read

An agent that answers questions about companies is only as reliable as what its tools return. This guide shows how to build the tool layer on structured company data: which lookups to offer, what every result should carry so that the model can date and cite what it says, and how to make the agent report that the data is silent instead of guessing. The examples use Fokals datasets, and the design holds for any dated, sourced company data.

Why rows serve an agent better than pages

A scraped page gives a model text of unknown age, written for a visitor, with the fact it needs somewhere inside it. A row gives the fact, its date and its source. Five properties of the Fokals datasets matter to an agent.

  • Dated. A technology in Technology Stack (company_technologies) has first_seen_at and last_seen_at, a change in Technology Changes has observed_at, an announcement in Company News has at. An answer that says as of 2 October can be checked.
  • Sourced. Every record names its source and says when it was observed, which gives the model a provenance to cite. Announcements and postings carry the url of the original.
  • Honestly empty. Labels are assigned by a model under named versions. A choice or score below 45 percent probability is left empty, and a flag is set only at 70 percent or more. An empty value means not known, and an agent can say so.
  • Compact. A company's hiring for one day is a row of counts and a few JSON maps. The same fact on a careers page sits among navigation, consent text and marketing copy.
  • Keyed. Every company has one stable Fokals company ID, and a listed company carries its ticker, MIC, ISIN, LEI and share-class FIGI, the company identifiers that tell two companies with the same name apart.

Pages remain the place for a company's own wording of what it does today. Rows give facts and change. For positioning, read the page, and treat what you read as a separate, undated source.

The tools to give the agent

Offer one tool per question, and make every tool except the first take company IDs. A name is never the key. Each tool reads one dataset, introduced here by its name and, beside it, its table.

ToolQuestionDataset and fieldsReturns
find_companyWhich company is this?company_id, company, ticker, mic, isin and figi on any company-level dataset, and domain in Technology Stack (company_technologies)Up to five candidates and the identifier that matched
get_stackWhat does its site run now?Technology Stack (company_technologies): technology_name, technology_category, first_seen_at, last_seen_at, missing_sinceCurrent rows with dates
get_changesWhat changed since a date?Technology Changes (company_tech_events): observed_at, category, key, change, before, afterRows since the date, oldest first
get_hiringHow is it hiring?Hiring Activity (company_hiring_daily): open_postings, new_postings, closed_postings, by_function, by_countryThe latest closed day and the trailing 28 days
get_announcementsWhat has it said about itself?Company News (company_news): at, title, url, event_types, itemsRows since a date, each with its link
get_intentWhy might it be buying?Intent Scores (company_intent_weekly): topic, score, surge, evidenceThe latest week's rows and their evidence

find_company accepts a domain, a ticker with its MIC, an ISIN, an LEI or a FIGI. It never picks silently when more than one company matches, and it returns not_covered when none does. The other tools take a since date, an as_of date and a limit, and nothing else the model could get wrong. The stack and change tools read the marketing stack dataset, and the announcement tool reads the announcements dataset.

Whether the tools call the REST API at request time or read tables you loaded from bulk files is your choice. Either way, cache the daily and weekly datasets by company and period. Those rows are final once their period has closed, so an entry never needs invalidating inside its period, and your rate limits by the minute and by the day are spent only on new periods. Expire entries from the event datasets at your next daily load.

What every result should carry

The model reads only what the tool returns, so put the date, the source and the limits in the result. The example is illustrative.

{
  "status": "ok",
  "company": "Acme Robotics",
  "company_id": "illustrative",
  "as_of": "2026-10-04T06:00:00Z",
  "source": "company_tech_events",
  "label_version": null,
  "rows": [
    {
      "cite": "c1",
      "observed_at": "2026-10-02T06:14:09Z",
      "category": "technology",
      "change": "added",
      "technology_name": "Salesforce",
      "technology_category": "CRM"
    }
  ],
  "note": "A site's first reading is a baseline. A change is something seen between two readings."
}

Five fields carry the discipline.

  • status is one of ok, no_rows, not_covered, ambiguous or out_of_scope. A model that cannot tell a quiet company from an unknown one will treat both as quiet.
  • as_of names the latest closed period behind the answer: the last closed UTC day for a daily dataset, the Monday of the week for a weekly one, and the time of the last reading of the site, last_read_at in Website Profile (company_site_facts), for website facts.
  • cite gives each row a handle that the model must use to refer to it.
  • label_version is set where a label is model output, so that the agent can say an event type was read by a model under news-v1.
  • note carries the dataset's reading rule in one sentence. The rule then travels with the answer instead of living in the prompt.

Telling the agent what a status means

Three statuses need three different answers.

Out of scope. Fokals is company-level data: technology, hiring, intent, announcements and scale for firms. A question such as who runs sales at Acme Robotics should return out_of_scope with a reason, because a model asked for a name will otherwise supply one from memory. A leadership change in Company News is delivered as a role, such as chief financial officer, which is the form the agent can cite. Put the list of question types the data answers in the tool descriptions, where the model reads it before it calls.

Coverage is explicit. The status not_covered separates a company your tools have no record for from a company that did nothing, so the agent never reads silence as a fact. The index spans equity listings in 79 countries, the brands they own and verified private companies.

No change seen. A detection shows presence on the website, and a site's first reading sets a baseline and creates no event. Phrase an empty result as no change between readings, and let the note field say so.

Citations from evidence

Every claim in an answer should resolve to a row. Announcements and postings carry the url of the original. An intent score carries evidence, whose entries give date, kind, source, weight and detail, so each score names the dated signals behind it and a model reviewer can open the source. A website fact is cited as its dataset and observation time. Render [c1] in the answer as a card that shows the dataset, the time and the link.

Check mechanically after the model replies. Every handle must exist in the tool results of that turn, and every number in the answer must appear in those results. A sentence that fails either check is removed or sent back for repair. Both checks are cheap to run and catch the commonest invention, a figure the tools never returned.

Keeping an agent inside its limits

A key carries scopes and rate limits by minute and by day. Keep it on your server, give the agent's key only the scopes it needs, and cap the tool calls per question, for example at eight. Let the tools accept a list of company IDs, so that a question about twenty companies is one call and not twenty.

Test the agent on a set built from the datasets, where the answers are computed and not written by hand.

  • Lookup. Resolve ten companies by domain, by ticker with MIC and by ISIN, and check each against the right company_id.
  • Change. Ask what a company added after 1 October and compare with Technology Changes (company_tech_events).
  • As of. Ask the same question as of an earlier date. Only rows observed by then may appear.
  • Scope. Questions about individuals must end in an answer that says what the data covers: companies, their technology, hiring, intent, announcements and scale.
  • Citation. Every claim has a valid handle and every number appears in a result.

Score each as pass or fail and report the rate by type. A drop in the as-of or scope tests is the signal to stop and fix.

What the agent receives

The agent gets facts and change, dated and sourced. A posting is delivered as structured fields, labelled by job function, seniority and ten role flags, and an announcement carries up to 1,200 characters of the company's own words with the link to the original. The data is refreshed daily to weekly, which suits answers about recent change, and every label comes from a named, frozen version, so the same question gives the same answer next month. The tool layer is yours to build over the REST API or bulk files, and the guide to serving company data to AI agents over Model Context Protocol shows how a team wraps an API as tools.

How Fokals delivers it

Every dataset named here is defined in the data dictionary, and the API reference lists the endpoints, scopes and limits. For the retrieval side of the same problem, see retrieval over structured company signals.

Frequently asked questions

How do I give an AI agent access to company data?

Expose a small set of read-only tools, one per question, over a REST API or datasets you have loaded: a lookup that resolves a domain, ticker or identifier to a company ID, and tools for its technology stack, changes, hiring, announcements and intent. Return dated rows with their source, a status that tells no result from no coverage, and a handle the model must cite.

Why use structured company data instead of web scraping for an agent?

Rows carry a date, a source and a stable company ID, while a scraped page carries text of unknown age with the fact buried in it. An agent can cite a row, check it against a date and say when a field is empty. A page remains the place for how a company describes itself today, so treat the two as different sources.

How do I stop an AI agent from inventing company facts?

Make the tools say when they have nothing: return a status of no rows, not covered or out of scope instead of an empty list. Require a citation handle for each claim, and check after the reply that every handle and every number appears in the tool results. Route questions about individuals to an out of scope answer that names what the data covers.

Can an agent explain why a company looks like a buyer?

Yes. Each intent score names its evidence: the five strongest dated signals behind it, each with its kind, topic and weight. The agent can cite the tool adoption, the opened role or the announcement that raised the score, and a reviewer can open the source. A surge flag marks a score that at least doubled against the company's own twelve-week average.

How fresh is the company data an agent receives?

Datasets are refreshed daily to weekly, depending on the dataset: the marketing stack daily to weekly, announcements daily to every three days, intent scores weekly with signals daily. Return an as-of date with every result and have the agent state it. The data suits answers about recent change in a company's technology, hiring and announcements.

How does an agent reach Fokals data?

Fokals is delivered direct, by a REST API of 25 endpoints and as bulk files in JSON, JSON Lines or CSV. You write the tool layer over the API or over datasets you have loaded, keep the key on your server, and give each client its own scopes and rate limits.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.