google3 min read

Curated summary

Towards developing future-ready skills with generative AI

Read original(opens in new tab)

Vantage is a Google Research experiment that uses generative AI to assess durable “future-ready” skills such as critical thinking, collaboration, conflict resolution, and creativity. It places students in realistic conversations with AI teammates, dynamically introduces challenges, and evaluates performance against educational rubrics. A study with New York University found that AI-generated scores agreed with human expert ratings at a comparable level to agreement between human raters.

Why Future-Ready Skills Are Difficult to Measure

  • Skills such as collaboration, creative thinking, and conflict resolution are increasingly important as technology changes work and education.
  • Traditional tests are too rigid to capture how people think, communicate, and respond in realistic situations.
  • Human-based assessments can be resource-intensive, difficult to standardize, and dependent on whether challenging situations arise naturally.
  • Vantage aims to make these skills measurable, scalable, and useful for guiding instruction and student growth.

AI-Simulated Team Assessments

  • Students participate in open-ended tasks, such as preparing a debate or pitching a creative idea, alongside AI avatars.
  • An “Executive LLM” uses an assessment rubric to manage the conversation and introduce targeted challenges, such as disagreement or conflict.
  • This adaptive process is designed to elicit enough evidence to assess a particular skill while keeping the interaction natural.
  • An “AI Evaluator” reviews the conversation transcript using the same rubric.
  • Students receive a visual skill map and qualitative feedback describing their demonstrated strengths and areas for improvement.

Validation with New York University

  • Google Research partnered with NYU to align Vantage’s tasks and scoring criteria with established educational rubrics.
  • The joint study involved 188 U.S. participants aged 18–25 and focused on conflict resolution and project management.
  • Researchers tested whether the Executive LLM could steer conversations toward specific skills.
  • Steered conversations produced significantly more skill-relevant information than conversations involving independent, uncoordinated AI avatars.
  • The AI Evaluator’s scores showed agreement with human expert ratings comparable to the agreement between two human raters.
  • The results suggest that LLM-based assessment can provide a scalable alternative for evaluating complex interpersonal skills.

Additional Research

  • Google also collaborated with OpenMic to study creativity and English language arts.
  • The collaboration analyzed work from 180 students completing creative multimedia assignments, including character interviews and literature-related media articles.
  • These studies tested whether the evaluation approach could extend beyond collaboration-focused tasks.

Vantage is available in English through Google Labs as a research experiment. Its approach could help educators provide more consistent practice, evidence-based feedback, and scalable assessment for skills that conventional tests struggle to capture.

Continue with another curated summary.