How accurate is it?
Results by age gap
- Original childhood photo, unchanged
- Aged by Still Looking
- Aged image vs. other people (chance level)
| Age gap | People | Original photo | Aged image | Other people | Right person first (original / aged) |
|---|---|---|---|---|---|
| Under 5 years | 14 | 0.52 | 0.23 | 0.04 | 93% / 57% |
| 5 to 10 years | 14 | 0.47 | 0.20 | 0.03 | 100% / 71% |
| Over 10 years | 14 | 0.27 | 0.10 | 0.02 | 64% / 14% |
| All | 42 | 0.42 | 0.18 | 0.03 | 86% / 48% |
What we did
- We used FG-NET, a research dataset with photos of the same people at different ages, and picked 42 people with a childhood photo (age 12 or under) and a later photo.
- We aged each childhood photo to the age in the later photo using exactly the same code as this website (same instructions, same AI model, 3 variations).
- A face-recognition model (ArcFace) scored how similar each image is to the person's real later photo. We compared that with doing nothing (the unchanged childhood photo) and with other people's photos.
Can it be improved?
Experiments on the same people, each changing only one thing and compared with our normal result.
More photos of the child
Up to 2 extra childhood photos, taken at the same age or younger, added next to the main photo.
Tested on 25 people. Similarity to the real later photo went from 0.19 to 0.16 (-0.03; 95% range -0.054 to +0.001), better for 40% of people. Right person picked first: 44% → 40%. The range includes zero, so we can't say it made a real difference.
Distinguishing features
An AI model listed visible lasting marks (moles, scars, eye colour) that were added to the instructions. In the app, families check this list first; in this test nobody did, so it is the worst case.
Tested on 13 people. Similarity to the real later photo went from 0.16 to 0.15 (-0.02; 95% range -0.050 to +0.014), better for 54% of people. Right person picked first: 46% → 46%. The range includes zero, so we can't say it made a real difference. (12 people had no clear features and were left out.)
Nano Banana 2 instead (now used by the app)
A newer image model from Google, given exactly the same instructions and photo as our free model. It is paid (about $0.07 per image) and has no random seed, so each run differs.
Tested on 42 people. Similarity to the real later photo went from 0.17 to 0.21 (+0.04; 95% range +0.003 to +0.070), better for 52% of people. Right person picked first: 43% → 48%. The whole range is above zero, so this helped.
An editing model instead (SAM)
Our model redraws a new face. SAM, a research face-aging model, edits the existing face instead and is trained to keep identity. It takes one photo only (no family photos or instructions).
Tested on 38 people. Similarity to the real later photo went from 0.18 to 0.20 (+0.03; 95% range -0.007 to +0.062), better for 66% of people. Right person picked first: 47% → 47%. The range includes zero, so we can't say it made a real difference.
Where it falls short
- The AI “beautifies” faces. Results tend to look smooth and symmetrical, like stock photos, which removes some of what makes a face unique.
- Longer gaps are much harder. Accuracy drops clearly when more than 10 years have passed.
- We could not check the age. The age-estimation model we tried guessed most children as adults, even in real photos, so we can't claim the images show exactly the right age.
- Family photos are not measured. No public dataset has a child's photos over time together with their parents' photos, so we can't yet say whether family photos help.
- Small, old dataset. FG-NET has 82 people, mostly scanned prints, and does not represent every ethnicity or age group. Results may differ for other children.
- A machine's view. Face-recognition similarity is a stand-in for “does this look like them?” We have not tested whether people recognise the aged images better.
Test run 2026-09-28 · image model: FLUX.2 [klein] 4B on Cloudflare Workers AI · face model: InsightFace buffalo_l (research use) · FG-NET photos are not shown or published here because their licence allows research use only.