int-bench-21ed28b5·1 events·first seen Aliases: Int-Bench
Researchers introduce Int-Bench, a simulation-based benchmark evaluating how LLMs intervene during student problem-solving across code debugging, mathematics, and brain teasers. The benchmark compares LLM 'teachers' to human tutors, finding that LLMs intervene more frequently, earlier, and tend to provide complete solutions rather than targeted hints. The study concludes that current LLM assistants optimize for immediate task success rather than fostering deeper reasoning and long-term learning, a meaningful finding for AI tutoring system design.