08/03/2026 • by Jonas Kellermeyer
Synthetic Data Analysis: Between the Promise of Insight and Methodological Caution
Few concepts today carry as much hope as that of synthetic data. Wherever real information is missing, difficult to access, or cannot be readily processed for good reason, artificially generated data suddenly appears as an elegant solution to pressing problems. But how well does synthetic data actually perform in everyday research practice? A balanced consideration.
Reflection as a Guarantee of Functionality
They promise scalability, are purportedly compatible with data protection requirements, invite experimentation, and moreover offer a certain degree of independence from the often laborious conditions of empirical data collection – in other words, synthetic data as a whole carries enormous expectations. This explicit promise weighs particularly heavily in fields where sensitive information is at stake, from health and education to questions of social participation. In this context, synthetic data analysis appears almost as a methodological royal road: one gains information without having to dirty one’s hands in the real world. Yet it is precisely at this point that the real work begins – namely, the work of reflection.
For synthetic data is – even less so than its conventionally collected counterparts — value-free or neutral (Gitelman 2013). Rather, it is the result of a modelling process in which a certain tendency in the weighting of data is already inscribed into the system at an early stage, usually on the basis of previously collected real-world data, though it is always worth inquiring into the chosen AI model. Anyone working with synthetically generated data is therefore never dealing merely with an equivalent substitute, but always with a necessarily distorted representation of what we commonly call “reality.” As Gitelman (2013) emphasizes, data are never “raw,” but always already processed; in the case of synthetic data, this prior mediation is effectively doubled by the generative model, meaning that existing social asymmetries are not dissolved, but instead risk being algorithmically reinforced (cf. D’Ignazio & Klein 2020).
What Are Synthetic Data, Exactly?
Synthetic data are not merely anonymized or cleaned-up real-world data. They are artificially generated data points, datasets, or trajectories designed to reproduce real structures as plausibly as possible without being identical to specific individuals or individual events (cf. Nikolenko 2021). Their value lies not in a high degree of realism, but rather in their often sharpened similarity to real-world data: they are meant to appear as if they had emerged from a particular reality – and it is precisely here that both their strength and their weakness reside.
For this similarity is never complete. Every process of synthetic data generation is based on selection: certain patterns are retained, others smoothed out; rare events are underrepresented, ambiguities reduced, contradictions silently resolved, and biases often greatly amplified. What remains is frequently a structured, highly legible, and model-compatible world. This can be helpful for technical development work. But when it comes to describing social reality and testing products accordingly, it becomes problematic rather quickly.
The Particular Appeal of Synthetic Data Analysis
There are several reasons why synthetic data currently appears so attractive: it allows systems to be prepared before sufficient real-world data is available. It can help test data pipelines, develop features, prepare visualizations, and formulate initial hypotheses. Above all, wherever real data collection is expensive, slow, or ethically sensitive, synthetic data offers a welcome alternative that can already be worked with. It makes possible what would otherwise only become conceivable much later: namely, working with data structures before the world has actually produced them.
This is a significant advantage, especially in early project phases. Anyone seeking to model baselines, routines, anomalies, or breaks does not have to wait until lengthy pilot phases have been completed. In this sense, synthetic data creates a preliminary space of insight. It helps sharpen questions without requiring immediately empirically robust answers.
Accordingly, one should not begin to confuse a synthetic data basis with empirical evidence. Synthetic data analysis can show how a system might function. It cannot show that the modeled reality has already been sufficiently understood, nor that the world actually corresponds to the model.
Where the Problem Arises
The real methodological difficulty of synthetic data analysis lies in the fact that it is mistakenly perceived as a pragmatic shortcut, when in truth it is better understood as a theoretically sharpening condensation. Every simulation carries within it a worldview. If, for example, household day-to-day trajectories are synthetically generated, decisions must first be made about what counts as a “typical day,” which deviations are considered relevant, which contexts are capable of explaining events, and which ruptures appear worthy of modelling. In other words: interpretation has already taken place before any data is synthesized.
This becomes particularly delicate when synthetic data is used to investigate explicitly social phenomena. For social realities are not only complex, but also ambiguous, and can register in highly subjective ways. The very same action can point in fundamentally different directions for different social actors. If, for example, a person leaves their home only rarely, this may be due to illness – but it may also indicate bad weather, voluntary withdrawal, visitors, fear, exhaustion, or the disappearance of a former routine. One and the same data structure can therefore contain several social truths at once. Synthetic datasets, however, tend to reduce this complexity because, by virtue of their modelling logic, they depend on a high degree of model-compatible consistency (cf. Esposito 2022).
It is precisely here that a dangerous confusion threatens to arise: what appears plausible in the simulation is all too quickly taken to be what is likely in reality. But plausibility is not yet proof or validation.
Synthetic Data as Structured Fiction
Perhaps it is more productive to understand synthetic data not as a substitute for reality, but rather as a form of structured fiction. They are useful because they are controllable. They can be varied, repeated, extended, disrupted, and systematically broken. One can use them to test what a system would “perceive” if certain routines appeared stable or if particular patterns changed over time. It is precisely here that their heuristic value lies: they do not so much grant access to truth as open up a corridor of possibilities.
Understood as structured fictions, synthetic data can help make blind spots in one’s own thinking visible. Which features do we consider relevant at all? Which breaks do we address, and which do we not? What notion of normality silently underlies our simulation? And where does the simulation reveal more about our own heightened expectations than about the target group itself? In this sense, synthetic data analysis can be highly illuminating – provided it remains inherently reflexive.
The Crucial Distinction: Careful Testing or Sheer Assertion?
Synthetic data analysis becomes methodologically sound only when it is clearly stated what it is meant to achieve – and what it is not. For development work, the use of synthetic data is particularly useful when it comes to
- preparing data models,
- testing feature logics,
- developing visualizations,
- simulating system reactions to breaks,
- and/or systematically working through edge cases under controlled conditions.
As soon as synthetic data is used to make robust claims about actually existing distributions, real risks, real acceptance, or even real-world effectiveness, its character changes. What was once a consciously heuristic methodology then turns into a form of pseudo-evidence. And it is precisely at this point that caution is required.
The decisive boundary, then, does not lie between the “real” and the “artificial,” but between development-supporting exploration and an inadmissible claim to validity. The benefits of synthetic data may and should be used in a preparatory way. To use them as a shortcut to validation, however, is more than merely foolish.
Why This Insight Matters Especially in Relation to Social Phenomena
Regarding purely technical domains, where signals tend to be clearer, the use of synthetic data may have less far-reaching consequences. In social contexts, however, the greatest caution is required. Anyone seeking to render routines, withdrawal, losses of participation, or even social isolation visible in data form should remain aware that these are not phenomena that simply exist as objective facts. They emerge through the interplay of behavior, context, interpretation, and relation. At best, synthetic data can offer approximations of observable patterns here – but it falls far short of capturing the meaning these patterns hold for real individuals.
This is not an argument against the use of synthetic data analysis. It is merely a reminder of the need for reflection – an argument against its uncritical application. Precisely in projects oriented toward the needs and risks of fragile and/or vulnerable groups, the distinction between data-based deviation and social reality must therefore be drawn all the more clearly. A technical system may recognize patterns and even replicate them in part in meaningful ways. But that still does not allow it to say what these patterns mean for a person’s life; the thoroughly human capacity for empathy cannot be outsourced.
The Actual Strength: Preparing Judgment
The perhaps most important contribution of synthetic data analysis therefore lies not in the production of artificial certainty, but in the cultivation of methodologically grounded judgment. Valid synthetic datasets compel us to make implicit assumptions explicitly visible. They reveal which notions of normality a project carries with it, which contexts ought to be taken into account, and where systems presume a false sense of clarity. In this respect, synthetic data is not a shortcut around empirical work, but much more a preparation for real-world testing.
This applies especially when synthetic data generation is understood not as a purely technical process, but as an interdisciplinary endeavor. Only where domain logic, technical modeling, and critical reflection come together can simulation become something more than an impressive speculation. The ability to ask better questions is the true game changer that synthetic data brings with it.
Conclusion
Synthetic data analysis is neither a miracle cure nor a mere placebo. Its value lies in the controlled production of what might be called spaces of possibility. Synthetic data can help prepare systems during development, generate hypotheses, and sharpen methodological decisions. Its weakness begins where synthetic data analysis presents itself as silent evidence and obscures the fact that it rests on sharpened model assumptions.
Anyone working with synthetic data should therefore not ask whether it is “real enough.” The more decisive question is: what, precisely, is it supposed to be sufficient for? For exploration, preparation, and reflection, it can be exceptionally valuable. For robust claims about social reality, however, it is explicitly insufficient. The fact that “[t]he body as representation relinquishes its sovereignty, leaving the image of the body available for appropriation and for reestablishment in sign networks separate from those of the given world” (Critical Art Ensemble 1994: 57) becomes particularly salient today in relation to the use of synthetically generated datasets. Yet this possibility also brings with it a responsibility that productively limits its application: synthetic data analysis is at its best when it does not pretend to be more than it actually is – a methodological tool that helps shape the world in a speculative manner without claiming any truly serious access to the lived world itself.
Sources
Critical Art Ensemble (1994): The Electronic Disturbance. Autonomedia, New York.
D’Ignazio, Catherine & Klein, Lauren F. (2020): Data Feminism. The MIT Press, Cambridge, Massachusetts.
Esposito, Elena (2022): Artificial Communication: How Algorithms Produce Social Intelligence. The MIT Press, Cambridge, Massachusetts.
Gitelman, Lisa (2013): „Raw Data“ Is an Oxymoron. The MIT Press, Cambridge, Massachusetts.
Nikolenko, Sergey I. (2021): Synthetic Data for Deep Learning. Springer Nature Switzerland, Cham.