trek-2010e478·1 events·first seen Aliases: TREK
TREK (Travel Reasoning and Evaluation Kit) is a new benchmark for evaluating tool-using LLM agents on multi-constraint travel itinerary synthesis, requiring plans to be simultaneously constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and persona-responsive. The benchmark comprises 800 tasks over a synthetic knowledge base of 212,530 records, scored by a fully deterministic rule-based evaluator with no LLM judge, making results reproducible and auditable. Evaluating 15 agents across nine constraint dimensions, even the strongest model (GPT-5.6) achieves only 46.2% success on solvable tasks, with satisfying unstated traveler needs emerging as the universal bottleneck. The dataset, tool sandbox, evaluator, and agent code are released publicly.