STRIPED DONKEYWe do the heavy lifting.
Spotlight

Synthetic City Twins · Part 2 of 2

Synthetic City Twins.
Evidence, assumptions and responsible use.

How to inspect synthetic city data: validation, assumptions, privacy and practical next steps for researchers and organizations.

Explore the Methodology behind the model

Connected histories are a starting point. Evidence determines what we can responsibly do with them.

Part 1 introduced Synthetic City Twins and followed a published employment history. Here we turn to validation, assumptions, privacy and the questions to ask before using a synthetic release.

Generating Data Is Easy. Keeping It Coherent Is Hard.

This is why certification is part of FCTE's construction process, not simply a quality badge added at the end.

Checks can operate at several levels.

Validation questions described in the article
Review areaQuestions to ask
Structural integrity

Do required relationships exist?

Are references valid?

Are entities connected correctly?

Temporal integrity

Does cause precede effect?

Do lifecycle events occur in possible sequences?

Does historical data accidentally contain future information?

Domain consistency

Does a record obey the rules of the domain being modeled?

Cross-domain consistency

If education, employment, healthcare, insurance or other domains refer to the same synthetic resident, do those histories remain logically compatible?

Empirical grounding

Where credible external evidence exists, how do selected synthetic distributions compare against it?

Release integrity

Can a user verify exactly which files and artifacts belong to a particular release?

A review framework for relationships, chronology, benchmarking and release integrity.

The purpose is not to claim perfection.

It is to make the system inspectable.

A dataset should not become trustworthy simply because its creator describes it as realistic.

Users should have something concrete to examine.

Transparency Matters More Than Pretending to Be Perfect

The word realistic is often used too casually in synthetic data.

A dataset can be internally coherent but poorly calibrated against an important external statistic.

It can resemble one reference population while being inappropriate for another use.

It can be structurally correct while simplifying a real-world phenomenon.

It can be excellent for software testing and unsuitable for estimating an official population rate.

Those distinctions matter.

A certification PASS therefore does not mean:

This is reality.

It means:

A defined test was executed and its stated acceptance criteria were met.

Likewise, a disclosed limitation is not automatically a weakness.

It defines the boundary of the claim being made.

For synthetic data to earn serious adoption — particularly in healthcare, research, financial services and public-interest applications — that discipline will matter more than grand claims of perfect realism.

Privacy Without Pretending Risk Disappears

Privacy is another reason synthetic data matters.

A genuinely synthetic population can support many forms of experimentation without providing analysts with the personal records of actual residents.

Developers can test data pipelines.

Researchers can experiment with models.

Students can work with complex relational datasets.

Teams can prototype applications before accessing restricted production systems.

This can substantially reduce dependence on personally identifiable information.

But synthetic data does not magically eliminate every risk.

Poor assumptions remain possible.

Bias remains possible.

Misinterpretation remains possible.

Models can still be used outside their intended scope.

The stronger promise is therefore not:

Zero risk.

It is:

Useful data environments with greatly reduced dependence on identifiable personal records, combined with explicit governance, validation and limitations.

That is a more durable foundation.

What Synthetic City Twins Can Enable

Connected synthetic populations create useful possibilities across several fields.

Machine Learning

Teams can build and evaluate classification, regression, survival, forecasting and event-sequence models using datasets with explicit temporal structure.

AI Evaluation

Synthetic populations can provide controlled scenarios for evaluating agents, copilots, retrieval systems and automated decision-support workflows.

Software Engineering

Applications can be tested against large populations without filling development systems with copied production records.

Longitudinal Research

Researchers can explore transitions, sequences and outcomes rather than being restricted to isolated cross-sectional tables.

Education

Students can learn data science, statistics, database engineering and machine learning using complex datasets without needing access to sensitive citizen records.

Scenario Exploration

Synthetic environments can support what-if analysis, provided simulated results are clearly separated from forecasts of the real world.

Data-System Development

Governments, companies and nonprofits can prototype architectures, reporting systems and analytical pipelines before exposing them to sensitive operational information.

Synthetic data does not need to replace real-world data to be valuable.

In many cases, it simply needs to make experimentation possible where experimentation was previously difficult.

Synthetic Data Does Not Replace African Evidence

This may be the most important point.

Africa still needs better real-world data.

Better surveys.

Better administrative systems.

More digitization.

Stronger interoperability.

Improved statistical capacity.

More locally generated evidence.

Better privacy infrastructure.

Synthetic City Twins do not remove those needs.

A synthetic healthcare dataset should not replace epidemiological surveillance.

A synthetic employment model should not replace a labour-force survey.

A simulated population should never quietly become an official population statistic.

Real-world evidence remains essential because simulation needs something against which its assumptions and outcomes can be challenged.

The relationship should therefore be complementary:

Real-world evidence helps constrain simulation.

Simulation enables experimentation where real records are unavailable, restricted or unsuitable.

Differences between simulation and observation reveal where the model needs improvement.

That feedback loop is more valuable than pretending one can permanently substitute for the other.

Africa Does Not Need to Wait

The opposite extreme should also be challenged.

Africa should not be told that meaningful AI development must wait until every hospital, school, employer, transport system, government registry and household has produced decades of perfectly digitized microdata.

That could take generations.

Researchers need usable datasets now.

Universities need teaching data now.

African ML engineers need benchmark problems now.

Startups need development environments now.

Software teams need realistic test populations now.

The alternative to imperfect real-world data cannot be no experimentation at all.

This is where carefully governed synthetic systems can become infrastructure.

Not because they manufacture truth.

But because they create environments in which teams can ask questions, build systems, discover failure modes and identify what better real-world evidence is still required.

The Real Leapfrog Opportunity

Africa has leapfrogged infrastructure constraints before.

Mobile phones expanded without universal fixed-line networks.

Mobile money expanded financial access beyond traditional bank branches.

Cloud computing allowed organizations to build systems without owning data centres.

Synthetic City Twins offer another possibility.

But the leapfrog is not:

Skip reality and generate everything.

It is:

Move from data scarcity to governed experimentation while simultaneously improving the real evidence needed to validate those experiments.

That creates a cycle:

Conceptual figure 03

Evidence improves the next iteration

  1. Evidence
  2. Simulation
  3. Testing
  4. Measurement
  5. Better Evidence
  6. Better Simulation

Each iteration improves the next.

A conceptual cycle: evidence shapes simulation, and measurement informs the next iteration.

Each iteration improves the next.

And that may ultimately be why regions with significant data constraints become some of the most important places for synthetic-data innovation.

The problem is difficult.

That is exactly why it is worth solving.

The Standard We Want to Build Toward

The future of African synthetic data should not be judged by how many billions of rows can be generated.

The better questions are:

Can we explain what was simulated?

Can we explain which evidence informed it?

Can we distinguish evidence from assumption?

Can we trace events through time?

Can we test relationships across domains?

Can independent users inspect the release?

Can we disclose what the model does not know?

Can the same synthetic resident remain logically coherent across a lifetime?

If the answer increasingly becomes yes, synthetic data becomes more than generated information.

It becomes research and development infrastructure.

A Different Way Forward

Africa has a data problem.

Synthetic data alone will not solve it.

But data scarcity should not mean the continent must wait on the sidelines of the AI era.

There is another path.

Build synthetic environments.

Ground them in the best evidence available.

Preserve relationships.

Respect chronology.

Certify what can be tested.

Benchmark what can be measured.

Disclose what remains uncertain.

Maintain a clear boundary between simulation and observation.

And improve the model as better evidence becomes available.

That is the path STRIPED DONKEY is pursuing with FCTE.

Nairobi and Lagos are the beginning.

More cities are coming.

More domains will connect.

More assumptions will be tested.

More evidence will challenge the models.

And the models will continue to evolve.

Not synthetic data that pretends to know Africa.

Synthetic infrastructure built to help Africa ask better questions.

STRIPED DONKEY
Connected synthetic data for AI, machine learning and research.

Powered by FCTE — Fidelra City Twin Engine.

Nairobi. Lagos. More cities coming.

← All Spotlight articles