Customers who build on your dataset load its history once and count on updates to keep their copy accurate. Later, they rerun old analyses and expect the old numbers. Financial data makes both hard, because releases are revised and corrected after they first appear, and occasionally withdrawn.

I build REST APIs for providers turning a dataset into that kind of product, starting with a written contract for what each record and timestamp means.

The contract comes before the endpoints

The contract starts from what customers do with the data, from following revisions to pulling a large range in bulk. It names the identifiers and units, says what each time field means and where it comes from, and shows how a revision, correction, or withdrawal appears to a customer. The endpoints are then designed to keep those promises.

Writing it takes access to the dataset and its sources, your entitlement rules, and time with the engineers who load and correct the data. Customer questions, especially about results they couldn’t reproduce, show which timestamps need the most care. We also agree on where the API runs and which of your systems it reads from.

The contract can’t promise more than your systems recorded. If nobody recorded when a revision became available to customers, that part of the history can’t be reconstructed, and the contract says so. A publication time copied from a source carries that source’s precision, and the contract states it.

Deliverables

  • The data contract and a data dictionary.
  • Query endpoints that read every page of a result from one snapshot.
  • Queries for the version available at a chosen time, and the revision history of each observation.
  • Bulk exports that hand off to a change stream without a gap.
  • Entitlement checks on every query, page, and download.
  • An OpenAPI specification, reference documentation, a local server with seeded data, and contract and authorization tests in CI.
  • Deployment configuration, and a written procedure for adding a field or dataset without breaking the contract.

Expected results are written first

Before writing the API, I work out from the source data what each query should return. That covers the version each cutoff selects, the records on every page, and the state a customer reaches after loading an export and applying the changes since. The contract tests check the API against those results.

The tests also change the data while customer jobs run, landing a revision partway through an export or delivering a page twice, and they expire cursors and revoke access between pages. The customer’s rebuilt copy still has to match the expected records, and a historical query still has to return what was available at its cutoff.

Changes after launch

I can stay on after launch for compatible changes, new endpoints and datasets, and operations. If a change would break existing customers, that work includes migrating them, with upgrade guides. SDKs and delivery monitoring are separate offers, and so is work on the pipeline that produces the data.

Message me on LinkedIn about your dataset and how customers get it today. If customers have found results they can’t reproduce, include those too.