Areas of Interest
Meta encourages innovative proposals that will generate novel, challenging, ground truth benchmarks and evaluations, both for pre-training and post-training. We are specifically seeking benchmarks with human-generated data (we will disqualify synthetically generated data unless it has been disclosed and the proposal has a detailed explanation justifying the quality and usefulness of it).
Knowing that there are more than 7,000 known languages in the world - and LLMs are just beginning to be available and trusted in a fraction of them - we want to understand how we can scalably extend the coverage for capabilities we care about into many languages . How can we make sure models in different languages align with human values in a culturally sensitive way? How can we measure capabilities that we care about? Proposals that can scalably extend to cover multiple languages or modalities will be prioritized.
We encourage submissions that utilize evaluations in the following areas:
Finalist: Antoine Bosselut, École Polytechnique Fédérale de Lausanne
MMRLU: Massive Multitask Regional Language Understanding
The INCLUDE project will reset standards for multilingual evaluation, developing benchmarks to assess LLMs in 100+ languages on region-specific knowledge and real-world reasoning that is rooted in the authentic settings where these languages are used
Finalist: Arman Cohan, Yale University
Evaluating and Advancing Complex, Real-world Reasoning in Large Language Models
This project focuses on systematically evaluating and improving large language models' capabilities in complex, logical reasoning, with an additional emphasis on legal reasoning and argumentation.
Finalist: Georgia Gkioxari, The California Institute of Technology
Modular Multi-modal Agents that Can [visually] Reason
This project introduces modular and adaptive frameworks powered by dynamic AI agents to advance the reasoning capabilities of vision and language models for fine-grained spatial understanding and 3D perception tasks
Finalist: Saikat Dutta, Cornell University
CodeArena: A Diverse Interactive Programming and Code Reviewing Environment for Evaluating Large Language Models
CodeArena is a next-generation benchmark for generative AI in software development, encompassing a diverse and realistic set of tasks, including bug fixing, responding to code reviews, test generation, multilingual code migration, and other essential workflows critical to modern software engineering.