A daily vendor feed asks four things of a pipeline: the files arrive, the table definition keeps up with them, each new day is processed once, and a bad day is caught before anyone queries it. This guide builds that pipeline in AWS Glue, which AWS describes as a serverless data integration service. It covers where the files land, what a crawler should and should not do, how a job reads only new partitions, when each step runs and which checks to attach. The tables and columns are those of the Fokals data dictionary, and the design suits any licensed feed in which each period is written once.
How the data reaches your account
Fokals is delivered by REST API and as bulk files in JSON, JSON Lines or CSV, with a manifest that names period, label versions and licence. Fokals is delivered direct, and you load it with the tools you already run in AWS. The first stage of the pipeline is a pull that lands the files in your bucket, ahead of Glue.
That stage is plain HTTPS. The API uses one bearer key per client, rate limits per key by minute and by day, and cursor pagination, and its feeds run oldest first from a time you set, so the last cursor you stored is your bookmark for incremental sync. Write the cursor only after the files are in S3, and keep the key in AWS Secrets Manager, which AWS describes as a service for managing, retrieving and rotating API keys and other secrets. The delivery page sets out the methods, formats and limits.
Run the code wherever you already run scheduled code. AWS's page on moving work off Glue Python shell jobs lists AWS Lambda, with a maximum timeout of 900 seconds, and Amazon ECS on AWS Fargate, with no maximum timeout. A long first pull fits a container task better than a function with that limit, and the daily increment can be either.
Lay out the bucket for the crawler
Give each Fokals dataset its own prefix and write each pull under a Hive-style key=value folder. For paths in that style a crawler takes the partition column name from the key and otherwise falls back to names such as partition_0, according to AWS's partition documentation. Name the key load_date, not day. The Athena CREATE TABLE reference says partition columns do not exist in the table data and that giving one the name of a table column is an error, and Hiring Activity (company_hiring_daily) from the hiring dataset already has a day column.
| Fokals dataset | Prefix under s3://acme-licensed/fokals/ | Key to test after loading |
|---|---|---|
| Hiring Activity | company_hiring_daily/load_date=2026-10-05/ | company_id and day |
| Sales Team Metrics | company_sales_weekly/load_date=2026-10-05/ | company_id and week_start |
| Market Series | market_series/load_date=2026-10-05/ | metric, dimension_kind, dimension, window_days and as_of |
| Job Postings | job_postings/load_date=2026-10-05/ | posting_id within one period |
The bucket and the buyer, Acme Robotics, are illustrative. Keep each manifest in the folder beside its files, so the sources, period, label versions and licence travel with the data.
Choose the file format with the crawler in mind. In the CSV files a list or object is JSON text in one cell. AWS's classifier page says the built-in CSV classifier accepts a first row as a header only when the rows beneath it parse as something other than strings, and otherwise names the columns col1, col2 and so on. It builds the table with LazySimpleSerDe and advises switching to OpenCSVSerDe, with inferred types set to string, when the data contains quoted strings, as a cell holding a JSON object does.
JSON Lines raises neither issue: there is no header row to detect and no quoting to unpick, and Glue's JSON reader treats each line as one record unless multiline is set. Land JSON Lines if you can choose. If you land CSV, define the table yourself from the dictionary.
Let the crawler describe the files, not rewrite them
A crawler classifies the files, groups them into tables or partitions and writes the metadata to the Data Catalog. Two of its behaviours matter for a licensed feed. It decides for itself when similar folders are partitions of one table, so AWS advises adding each table's root folder as a separate include path. And by default it updates the table to match whatever it finds. For a feed whose columns are documented you usually want to be told about a new column, not to absorb it silently. Fokals announces a breaking change at least 90 days ahead under a new version name, so a schema change is something to schedule.
The setting that fits is Crawl new sub-folders only. After a first full crawl, later runs add only new partitions. AWS says this mode does not detect changes to existing partitions, which suits tables written once, and that it forces the update and delete behaviours to LOG: a folder with an incompatible schema is not added, and the detail goes to CloudWatch Logs. Treat that log entry as an alert.
aws glue update-crawler \
--name fokals-raw \
--recrawl-policy RecrawlBehavior=CRAWL_NEW_FOLDERS_ONLY \
--schema-change-policy UpdateBehavior=LOG,DeleteBehavior=LOGTime-based schedules in Glue use cron syntax in UTC with a minimum precision of five minutes, according to AWS's schedule page, and cron(15 12 * * ? *) means every day at 12:15 UTC. UTC is also the clock of the feed: day is a closed UTC day and week_start is the Monday of a Monday-to-Sunday UTC week, so nothing needs converting.
Read only the new days
A Glue Spark job reads the catalog table, converts types and writes Parquet. Fokals delivers CSV, JSON and JSON Lines, and the job converts them to the Parquet files your lake reads. Job bookmarks keep the state that tells a rerun which input it has already processed. For an S3 source Glue compares the last modified time of the objects, it tracks a source only where the script passes transformation_ctx, and the job parameter --job-bookmark-option must be set to job-bookmark-enable. Python shell jobs cannot use bookmarks.
import sys
from awsglue.context import GlueContext
from awsglue.job import Job
from awsglue.utils import getResolvedOptions
from pyspark.context import SparkContext
from pyspark.sql import functions as F
args = getResolvedOptions(sys.argv, ["JOB_NAME", "from_date"])
glue = GlueContext(SparkContext())
glue.spark_session.conf.set("spark.sql.sources.partitionOverwriteMode", "dynamic")
job = Job(glue)
job.init(args["JOB_NAME"], args)
raw = glue.create_dynamic_frame.from_catalog(
database="fokals_raw",
table_name="company_hiring_daily",
transformation_ctx="hiring_in", # the bookmark key: never rename it
push_down_predicate=f"load_date >= '{args['from_date']}'",
)
typed = (
raw.toDF()
.withColumn("day", F.to_date("day"))
.withColumn("open_postings", F.col("open_postings").cast("int"))
.withColumn("new_postings", F.col("new_postings").cast("int"))
.withColumn("closed_postings", F.col("closed_postings").cast("int"))
)
(typed.write.mode("overwrite")
.partitionBy("load_date")
.parquet("s3://acme-licensed/curated/company_hiring_daily/"))
job.commit() # saves the bookmark state; without it nothing is rememberedThree cautions apply. A bookmark lists every file under each input partition, and the pushdown predicate narrows which partitions are listed, so set from_date earlier than the longest outage you could have, or a day missed in an outage is never read. Bookmarks track sources and not targets, so a run that wrote output and failed before job.commit() is repeated on the same input; the dynamic partition overwrite above replaces the partition instead of doubling it. And objects modified since the last run are processed again, so never rewrite a landed file in place: land a correction as a new object.
Order the steps and set the schedule
Glue has three trigger types: scheduled, conditional and on demand. A conditional trigger watches named jobs or crawlers for a status such as succeeded and starts the next. AWS's trigger page adds two constraints: one trigger can activate only two crawlers, and every job or crawler in a chain must descend from a single scheduled or on-demand trigger. For more than a couple of steps AWS prefers workflows, which can run on demand, on a schedule or from an Amazon EventBridge event.
The chain for the daily feed is a crawler for the raw prefix, the job above, and a crawler for the curated prefix. A daily table cannot exist before its day has closed, and a weekly table cannot exist before Sunday ends, so the pull step checks the manifest for the period and ends without landing anything when the period is not there yet. Start the workflow from the event that step sends when the files are in the bucket, or from a cron schedule with margin after UTC midnight and let the first step fail loudly on a missing period.
Check every load
AWS Glue Data Quality evaluates rules written in the DQDL language against Data Catalog tables or inside jobs. AWS describes the job form as proactive: it runs before data is loaded into the lake and can identify the exact rows that failed. For Hiring Activity, whose rows are per company and closed UTC day:
Rules = [
RowCount > 0,
IsComplete "company_id",
IsPrimaryKey "company_id" "day",
ColumnValues "open_postings" >= 0,
ColumnValues "new_postings" >= 0,
ColumnValues "closed_postings" >= 0
]The two-column IsPrimaryKey catches a pull that landed the same day twice. AWS's release notes say ColumnValues does not let a null pass a comparison, so use it only on columns that must be filled. For Job Postings, a rule such as ColumnValues "label_version" in ["jobs-v1", "jobs-v2"] flags a version you have not met. Fokals gives at least 90 days' notice of a breaking change, so a failure is a prompt to read the notice and update the list.
Count the rows flagged as reconstructed as a separate measure. Fokals writes each daily and weekly period once and marks a period written more than seven days after it closed with reconstructed=true: the rows are valid but late, and any test of what was known on a date should take the flag into account. The simplest completeness measure is the number of curated partitions against the number of closed days since your first pull. A gap is a period that never reached the bucket.
Where this stops
Glue catalogues, schedules and transforms; it does not decide what a company is. Joining to your security master on ISIN or FIGI, answering point-in-time questions and keeping the licence terms in view are your modelling and governance work. For access to catalog tables, AWS lists Lake Formation as the authorization layer for the Data Catalog. Every Fokals observation is dated and every period is written once, so a back-test reads the record as it stood on each day.
If you only need to query the files where they lie, an external table in Athena over the prefix needs no job at all, and the guide to querying company data on S3 with Athena shows that route.
Frequently asked questions
How do I load only new files with AWS Glue?
Turn on job bookmarks by setting --job-bookmark-option to job-bookmark-enable, read through the Data Catalog with the transformation context set, and call job.init at the start and job.commit at the end. For an S3 source Glue filters on the last modified time of objects, so a file modified after a run is processed again. Add a pushdown predicate on the partition key to limit which partitions are listed.
How often should a Glue crawler run for a daily feed?
Once per delivered period, after the files land. Set the crawler to crawl new sub-folders only, so each run adds just the new partition, and start it from a workflow that follows your pull step or from a UTC cron schedule. A crawler cannot tell a late period from a missing one, so compare the partitions it has added with the period named in the manifest.
Should I use a Glue crawler or define the table myself?
Define the table yourself when the columns are documented, as the Fokals data dictionary documents them, and you want fixed types. A crawler infers from the files: its CSV classifier decides about a header by heuristics, and AWS advises changing the SerDe for quoted strings. If you keep a crawler, set its schema change policy to log changes so it adds partitions and leaves the columns alone.
Where should the API pull run when the rest of the pipeline is in AWS Glue?
Anywhere that can make scheduled HTTPS calls and write to S3. AWS lists Lambda and Amazon ECS on Fargate as places to run scripts that once ran as Glue Python shell jobs, and gives Lambda a maximum timeout of 900 seconds. Keep the bearer key in a secrets store, store the last cursor only after the files are in S3, and let Glue take over from the bucket.
How does Fokals data get into S3 and AWS Glue?
Fokals is delivered direct, by REST API and as bulk files in JSON, JSON Lines or CSV. You call the API or download an export, land the files in your own bucket, and Glue takes over from there. The manifest that comes with an export names its period, label versions and licence, so each landed folder carries its own record.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.
What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.