Benchmarks
Coming
Coming
Soon
Learning from the Flip Zone demo and sharing what we discover.
The questions
We start by asking questions like:
- How can you define fun?
- Can an AI model predict whether humans will find something fun or challenging?
- Can an AI model independently create something fun and challenging? If not, what is stopping it?
- Can AI models play the game better than humans? Where are the gaps?
- Why does Model A perform Task 1 better than Model B, while Model A cannot do Task 2 at all and Model B struggles through it?
- How do these capabilities change over time?
As we seek answers to these questions and more, we aim to discover what patterns emerge over time.
See Flip Zone