Skip to content
Bits&Chips
×
×
About
Memberships
Advertising
Magazines
Videos
Contact

Log in

Jan Bosch is a research center director, professor, consultant and angel investor in startups. You can contact him at jan@janbosch.com.

Opinion

Who needs data when you can create it?

13 April 2026
Reading time: 5 minutes

The most interesting organizations going forward will use real-world data to ground their models in reality, but rely on synthetic data and simulation to scale, explore edge cases and accelerate learning.

Over the last decade, many of the companies I work with through Software Center have made significant investments in data. Sensors have been deployed, systems instrumented and pipelines built to collect and store vast amounts of information. In principle, this should provide a strong foundation for data- and AI-driven innovation.

In practice, however, a recurring pattern emerges. When a team wants to develop a specific use case, such as predictive maintenance for a rare failure mode, perception for an edge case in autonomous driving or a new optimization algorithm, they often discover that the required data is either not available, too sparse, biased or simply inaccessible due to regulatory constraints. GDPR, regional data residency requirements, contractual limitations and internal governance policies frequently prevent data from being used in the way it was originally intended. As a result, despite “having a lot of data,” companies often lack the right data.

This is where synthetic data and simulation enter the picture. Rather than relying exclusively on real-world data, companies can generate artificial data that mimics reality. Synthetic data allows organizations to create training environments where edge cases can be produced on demand, sensitive information can be removed by design and scenarios can be explored that would be prohibitively expensive, or even impossible, to capture in the physical world.

Startups such as Gretel.ai (bought by Nvidia last year) and Parallel Domain are building platforms that enable exactly this. They provide tools to generate realistic images, sensor data, tabular datasets and entire simulated environments tailored to specific use cases.

The applications are already compelling. In autonomous vehicles, synthetic data is used to train perception systems on rare but critical scenarios, such as unusual weather conditions or near-accidents, that would take years to observe in the real world. In robotics, simulation environments allow systems to learn tasks through millions of iterations without the cost and wear of physical hardware. In industrial settings, digital twins replicate factories or products, enabling experimentation and optimization without disrupting operations. And for many data-driven applications, synthetic datasets provide a way to develop and test models without exposing sensitive personal or proprietary information.

Beyond the technical capabilities, synthetic data introduces two important strategic shifts. First, for incumbents, it provides a way to circumvent the constraints of real-world data. Instead of being limited by what has been collected and what’s legally permissible to use, organizations can generate the data they need, aligned with their use case and compliant by design. This is particularly important in regulated industries, where the friction associated with data access is often one of the main bottlenecks to innovation.

Second, for startups, synthetic data changes the competitive landscape. Traditionally, access to large proprietary datasets has been a key barrier to entry. Companies with scale had a significant advantage simply because they owned more data. Synthetic data weakens this advantage. Startups can now create high-quality training data without owning massive real-world datasets, allowing them to compete more effectively with incumbents.

Synthetic data comes with important limitations that are easy to underestimate

Despite its promise, synthetic data comes with important limitations that are easy to underestimate. The central challenge is one of fidelity: How do we know that the generated data truly reflects the real world? Even small deviations in statistical distributions, correlations or edge case frequencies can lead to models that perform well in simulation but fail in practice. This is often referred to as the “sim-to-real gap,” particularly in domains such as robotics and autonomous systems. If the synthetic environment doesn’t capture the complexity, noise and unpredictability of reality, systems trained on it may learn the wrong abstractions. In addition, synthetic data generation itself encodes assumptions about what matters, what can be ignored and how variables interact, which may introduce hidden biases.

As a result, synthetic data rarely replaces real-world data entirely. Instead, it needs to be continuously validated and calibrated against real observations, with careful testing to ensure that models trained in simulated environments generalize reliably when deployed in the real world.

Leading companies in synthetic data and simulation have converged on a set of practical tactics to make artificial data useful in the real world. One of the most important is domain randomization: Instead of trying to perfectly replicate reality, they deliberately vary parameters such as lighting, textures, object positions and sensor noise across a wide range so that models learn to generalize rather than overfit to a single ‘perfect’ simulation. Closely related is the practice of iterative sim-to-real validation, where models trained on synthetic data are continuously tested and fine-tuned on small amounts of real-world data to reduce the so-called reality gap. Another key tactic is increasing fidelity where it matters, such as using high-quality 3D assets, physically based rendering and realistic sensor models, to close appearance and content gaps between simulation and reality. At the same time, leading players increasingly incorporate domain knowledge into the generation process, ensuring that synthetic scenarios reflect real-world constraints and edge cases rather than purely random variation. Finally, many combine synthetic and real data in a hybrid approach: Large-scale synthetic data is used for pre-training, while smaller real datasets are used for calibration and validation. Together, these tactics reflect a shift from simply generating data to engineering data generation as a core capability, where the goal isn’t realism per se, but reliable transfer of learning into the real world.

The whole picture points to a broader shift in how we think about data and advantage. Historically, success in data-driven systems was closely tied to owning the world, as in instrumenting reality, collecting data at scale and building proprietary datasets. Increasingly, however, we see the emergence of an alternative approach: simulating the world. Instead of waiting for data to be generated, companies can create it. The most interesting organizations going forward will likely combine both approaches. They’ll use real-world data to ground their models in reality, but rely on synthetic data and simulation to scale, explore edge cases and accelerate learning. In that sense, synthetic data isn’t just a technical solution; it represents a shift from a passive to an active approach to data: from collecting what happens to creating what’s needed. To return to Peter Drucker: “The best way to predict the future is to create it.”

Related content

The shape of what’s coming: a synthesis

You can’t address AI’s risks by banning ASML’s exports

Top jobs
Your vacancy here?
View the possibilities
in the media kit
Events
Courses
Headlines
  • ASIC design team spins out from Philips as ICwaves

    8 July 2026
  • Dutch defense embraces Intelic’s software-first drone interoperability approach

    8 July 2026
  • TNO and Destinus collaborate on radar seekers

    8 July 2026
  • Noviotech Campus sees another director go

    6 July 2026
  • UT appoints Gregor Halff as new executive board president

    6 July 2026
  • Reports: Semi equipment market to reach up to $250B in 2028

    1 July 2026
  • Japanese researcher proposes simpler high-NA EUV optics design

    25 June 2026
  • ASML partners with TNO on photonic chip pilot line

    24 June 2026
  • Dutch minister pushes back on US bid to tighten China chip controls

    24 June 2026
  • Alixlabs launches beta APS platform for non-litho patterning

    23 June 2026
  • Thales NL expands radar system production and test capacity

    23 June 2026
  • Robin Radar and TNO develop airborne radar

    22 June 2026
  • Nearfield closes record 330-million-euro funding round

    22 June 2026
  • Koen ten Hove appointed as CTO of Thales NL

    19 June 2026
  • Besi raises long-term targets on AI packaging and hybrid bonding demand

    18 June 2026
  • ABN Amro: AI and defense boom shakes up Dutch EMS sector

    17 June 2026
  • Nexperia stays profitable despite China disruption

    16 June 2026
  • ASML, TSMC and Imec scale 2D transistors to 50nm pitch on 300mm wafers

    15 June 2026
  • Optical interconnect market explodes to $39B by 2030

    15 June 2026
  • Forced layoffs avoided until May 2027 in union-backed ASML restructuring plan

    11 June 2026
Bits&Chips logo

Bits&Chips strengthens the high tech ecosystem in the Netherlands and Belgium and makes it healthier by supplying independent knowledge and information.

Bits&Chips focuses on news and trends in embedded systems, electronics, mechatronics and semiconductors. Our coverage revolves around the influence of technology.

Advertising
Subscribe
Events
Contact
Follow us on
High-Tech Systems Magazine (Dutch)
(c) Techwatch bv. All rights reserved. Techwatch reserves the rights to all information on this website (texts, images, videos, sounds), unless otherwise stated.
  • About
  • Memberships
  • Advertising
  • Videos
  • Contact
  • Search
Privacy settings

Bits&Chips uses technologies such as functional and analytical cookies to improve the user experience of the website. By consenting to the use of these technologies, we may capture (personal) data, unique identifiers, device and browser data, IP addresses, location data and browsing behavior. Want to know more about how we use your data? Please read our privacy statement.

 

Give permission or set your own preferences

Functional Always active
Functional cookies are necessary for the website to function properly. It is therefore not possible to reject or disable them.
Voorkeuren
De technische opslag of toegang is noodzakelijk voor het legitieme doel voorkeuren op te slaan die niet door de abonnee of gebruiker zijn aangevraagd.
Statistics
Analytical cookies are used to store statistical data. This data is stored and analyzed anonymously to map the use of the website. De technische opslag of toegang die uitsluitend wordt gebruikt voor anonieme statistische doeleinden. Zonder dagvaarding, vrijwillige naleving door je Internet Service Provider, of aanvullende gegevens van een derde partij, kan informatie die alleen voor dit doel wordt opgeslagen of opgehaald gewoonlijk niet worden gebruikt om je te identificeren.
Marketing
Technical storage or access is necessary to create user profiles for sending advertising or to track the user on a site or across sites for similar marketing purposes.
  • Manage options
  • Manage services
  • Manage {vendor_count} vendors
  • Read more about these purposes
View preferences
  • {title}
  • {title}
  • {title}

Your cart (items: 0)

Products in cart

Product Details Total
Subtotal €0.00
Taxes and discounts calculated at checkout.
View my cart
Go to checkout

Your cart is currently empty!

Start shopping

Notifications