Background: Many machine-learning tools can predict a child's future weight status, but they act as "black boxes" — clinicians cannot see why a prediction was made, which limits trust and clinical use. We used a new approach called DeepPySR, which searches for plain mathematical equations that predict BMI as accurately as black-box models, but in a form a clinician can read, check, and explain to a family. We tested this in the Raine Study (Gen 2), an Australian birth cohort followed from birth to young adulthood.
Methods: We combined genetic risk scores for BMI with early-life clinical and family information (e.g., birth weight, maternal factors, growth measurements) to build three types of models spanning ages 8 to 27 years: (1) models predicting BMI at one age at a time; (2) a single equation predicting BMI at any age; and (3) forecasting models that predict a child's future BMI using their own earlier growth history plus genetic risk. Each model's accuracy was compared against standard prediction methods, including regression, tree-based models, and neural networks.
Results: Our approach was consistently more accurate than the best standard method across all ages studied. For example, at age 23, our equation explained 60% of the variation in BMI versus 28% for the best standard method. A single equation predicting BMI across the whole age range performed comparably to standard methods at most ages. When forecasting future BMI from a child's own growth trajectory and genetic risk, our approach was again the most accurate on average across ages 8 to 27, while still producing a short, readable formula.
Conclusion: It is possible to predict BMI trajectories using compact, human-readable equations without sacrificing accuracy compared to complex models. They offer clinicians a transparent tool for identifying children and adolescents at risk of unhealthy weight gain, and for explaining that risk to families in understandable terms.