Beyond Student Labels
Does AI generate different-quality practice problems depending on how we describe the student's level, language background, or the prompt wording?
Method
We generate practice problems for the same set of topics under six prompt conditions, then have raters score each problem on an 8-dimension rubric (correctness, clarity, difficulty fit, language burden, scaffolding, tutor usefulness, cultural accessibility, and topic alignment). The platform stores 123 rater-level scores across 80 problems and exports the dataset for analysis.
Early finding
The Scaffolded condition currently scores highest overall (4.4/5), versus 3.7/5 for the unguided Basic baseline. Language-aware prompts gain the most on clarity and cultural accessibility - the dimensions that matter most for multilingual learners.