"How much data do we need and how long will it take to collect?" comes up on every call, usually with a worried glance at the calendar. The honest answer is "less than you fear and more carefully than you think." Data quantity is rarely the bottleneck in industrial sensing projects. Data quality, label quality and representativeness are. This is how we approach dataset engineering, and how to estimate what your project needs.

Three questions that decide the data budget

  1. How rare is the event? Detecting a defect that occurs once per thousand parts needs a strategy for capturing enough positives: historical scrap, deliberate provocation on a test rig, or augmentation grounded in physics. Continuous condition tracking needs less: the signal is always there.
  2. How variable is the process? Every product variant, recipe, shift and season the model must survive should be represented. Two weeks across the full product mix beat two months of a single product.
  3. How good are the labels? A label is a claim about what happened. If inspection results, timestamps and sensor recordings do not line up, the dataset is smaller than it looks. Aligning them is often the real work.

For a typical feasibility study, a few hours of recordings with known outcomes is enough to answer whether the signature exists. For a proof of concept, days to a few weeks of collection across the process window is typical. Multi-month campaigns are the exception, usually for seasonal effects or very rare failures.

What a dataset needs beyond samples

A label specification

Written down, with examples and edge cases, agreed by the people who will label. "Defective" is not a label. "Porosity visible in CT scan above 0.3 mm" is.

Metadata

Machine, product, recipe, shift, sensor position, ambient conditions. Without metadata you cannot diagnose why the model fails on Tuesday nights.

Quality assurance

Sample checks, inter-annotator agreement where humans label, automated checks for clipping, dropouts and clock drift in the recordings.

Versioning

Every model is trained on a specific dataset version. When the line changes, you add a version, you do not overwrite. This is how the dataset stays useful for the next project and the next vendor.

Why the dataset is the real asset

Models age. Sensors get replaced, products change, better architectures appear. A well-engineered dataset with clean labels and metadata is reusable across all of that. It is also the one thing a plant can own outright, independent of any vendor. That is why every project we run delivers the dataset, its specification and its documentation to the client, and why we offer dataset creation as a service in its own right, including for companies that want to train models with their own team.

Ask a vendor for the model and you get a file. Ask for the dataset and you get an asset.

Open benchmarks and why we contribute to them

We co-authored IDMT-Traffic, an open benchmark dataset for acoustic vehicle detection under realistic microphone mismatch. Public benchmarks with honest evaluation splits are how the field learns what actually works, and how clients can check claims instead of trusting them. Where a client's data cannot be shared, we design internal benchmarks with the same discipline: fixed splits, strong baselines, realistic test conditions.

Key takeaways

  • Quantity is rarely the bottleneck; representativeness and label quality are.
  • Rarity, variability and label quality decide the data budget.
  • A feasibility study needs hours; a proof of concept needs days to weeks across the process window.
  • Label specification, metadata, QA and versioning turn recordings into a dataset.
  • The dataset outlives the model; make sure you own it.

If any of this matches a problem on your line, the fastest way to find out what is possible is a free discovery call followed, where it makes sense, by a feasibility study of two to ten days.

A
AcousticAI Lab engineering teamIndustrial acoustic, vibration and multi-modal sensor AI · Edge and sovereign deployment · Germany and India