Curated summary
Towards developing future-ready skills with generative AI
Vantage is a Google Research experiment that uses generative AI to assess durable “future-ready” skills such as critical thinking, collaboration, conflict resolution, and creativity. It places students in realistic conversations with AI teammates, dynamically introduces challenges, and evaluates performance against educational rubrics. A study with New York University found that AI-generated scores agreed with human expert ratings at a comparable level to agreement between human raters.
Why Future-Ready Skills Are Difficult to Measure
- Skills such as collaboration, creative thinking, and conflict resolution are increasingly important as technology changes work and education.
- Traditional tests are too rigid to capture how people think, communicate, and respond in realistic situations.
- Human-based assessments can be resource-intensive, difficult to standardize, and dependent on whether challenging situations arise naturally.
- Vantage aims to make these skills measurable, scalable, and useful for guiding instruction and student growth.
AI-Simulated Team Assessments
- Students participate in open-ended tasks, such as preparing a debate or pitching a creative idea, alongside AI avatars.
- An “Executive LLM” uses an assessment rubric to manage the conversation and introduce targeted challenges, such as disagreement or conflict.
- This adaptive process is designed to elicit enough evidence to assess a particular skill while keeping the interaction natural.
- An “AI Evaluator” reviews the conversation transcript using the same rubric.
- Students receive a visual skill map and qualitative feedback describing their demonstrated strengths and areas for improvement.
Validation with New York University
- Google Research partnered with NYU to align Vantage’s tasks and scoring criteria with established educational rubrics.
- The joint study involved 188 U.S. participants aged 18–25 and focused on conflict resolution and project management.
- Researchers tested whether the Executive LLM could steer conversations toward specific skills.
- Steered conversations produced significantly more skill-relevant information than conversations involving independent, uncoordinated AI avatars.
- The AI Evaluator’s scores showed agreement with human expert ratings comparable to the agreement between two human raters.
- The results suggest that LLM-based assessment can provide a scalable alternative for evaluating complex interpersonal skills.
Additional Research
- Google also collaborated with OpenMic to study creativity and English language arts.
- The collaboration analyzed work from 180 students completing creative multimedia assignments, including character interviews and literature-related media articles.
- These studies tested whether the evaluation approach could extend beyond collaboration-focused tasks.
Vantage is available in English through Google Labs as a research experiment. Its approach could help educators provide more consistent practice, evidence-based feedback, and scalable assessment for skills that conventional tests struggle to capture.
Related reading
Continue with another curated summary.