Skip to content
Research

Beyond Student Labels

Does AI generate different-quality practice problems depending on how we describe the student's level, language background, or the prompt wording?

Method

We generate practice problems for the same set of topics under six prompt conditions, then have raters score each problem on an 8-dimension rubric (correctness, clarity, difficulty fit, language burden, scaffolding, tutor usefulness, cultural accessibility, and topic alignment). The platform stores 123 rater-level scores across 80 problems and exports the dataset for analysis.

Early finding

The Scaffolded condition currently scores highest overall (4.4/5), versus 3.7/5 for the unguided Basic baseline. Language-aware prompts gain the most on clarity and cultural accessibility - the dimensions that matter most for multilingual learners.

Scaffolded
4.4
Multilingual
4.2
Simple English
4.2
Culturally accessible
4
Grade-level
3.9
Basic
3.7
Want the full methodology and dataset? Explore the live research workspace →