When a team keeps third-party data for research, someone eventually asks whether the stored dataset is complete. The answer has two parts. One is whether every record the pipeline accepted reached storage. The other is which periods the source never delivered. I build pipelines that collect API data into storage you own and keep both parts of that answer with the data.

For a live feed, the pipeline holds incoming events durably and writes them to object storage in batches. For a query API, it keeps its place through long downloads and follows updates afterward.

Your queries decide the storage design

The way you plan to query or replay the data decides the schema, the file format and partitioning, how long data is retained, and how quickly new records must arrive. A first project covers one source and one destination. It can use an existing client or include building one.

You provide access to the source, a destination account or environment, and your limits on retention and infrastructure cost. We work out the expected volume, the failures the pipeline must tolerate, and how long collection should continue while storage is unavailable.

Records become durable before they wait for a batch

The pipeline has a defined point at which a record counts as accepted, and accepted records survive while they wait to become part of a file. A batch leaves the pending queue only after its file and its completion record are both confirmed. Retries keep each record’s identity, so a repeated delivery creates no duplicate.

Batch size and maximum wait trade file efficiency against freshness. Where the research needs it, the original payloads and receipt times are stored alongside the queryable form for later inspection or reprocessing.

Which failures are in scope decides how much protection the pipeline needs. Surviving a process restart, the loss of a host, and a long storage outage each need something different. Capacity and retention limits are part of the design, with alerts that warn your team before they’re reached.

Storage can’t recover what the source never sent

Durability covers records after acceptance. Events emitted while the collector was disconnected depend on whether the source can replay them, and a fresh snapshot restores current state without the history in between. The pipeline records known interruptions, and where completeness can’t be established, the dataset says so.

Testing recovery against known records

The tests feed the pipeline a prepared set of records, interrupt collection and delivery at chosen points, restore them, and compare what’s stored with what was sent. The cases include a process restart, a storage outage, and a write that succeeds while its response is lost. I also check batch age, readable output, and duplicate handling at the expected volume.

You receive the collector and storage code with its deployment configuration, a data dictionary, and examples that query or replay the stored data. The operating instructions cover backlog, failed writes, capture interruptions, and recovery. The test results and recovery commands come with them, along with any source behavior that remains unresolved.

I can also run or maintain the pipeline after handover, including monitoring, dependency updates, schema changes, capacity reviews, and new destinations. Diagnosing or fixing a pipeline you already run is a different project.

Start a conversation on LinkedIn with the source, where the data should live, and roughly how much arrives and how fresh it needs to be. It also helps to know what losing a day of data would stop your team from doing.