Andrew Dai

Andrew Dai

Co-founder & CEO, Elorian AI

Andrew Dai spent 12 years as a Research Scientist at Google Brain and DeepMind. He wrote the 2015 paper that OpenAI later cited as the original recipe for ChatGPT, was a core Lead on Gemini, GLaM, and PaLM 2, and his published research has accumulated over 67,000 citations. Now, he leads Elorian, a company building AI systems that understand the visual medium and apply reasoning the way humans do. Elorian recently launched with $55M at a $300M valuation, backed by Menlo Ventures, Altimeter, Striker Venture Partners, NVIDIA and Jeff Dean.

All Sessions by Andrew Dai

October 21, 2026

The Best Models Still Reason Visually Like Toddlers
2:00 pm - 2:45 pm
Room 203

Frontier AI models score 80–90% on standard benchmarks like RKGI, yet when tested on visual tasks any 3-year-old handles effortlessly (like counting objects in an image), those same models fall to pieces. Andrew watched this gap widen firsthand during his 14 years at Google Brain and DeepMind, where he co-led development on GLaM, PaLM 2, and Gemini. The problem is that most models hit high RKGI scores not through genuine visual understanding, but by coding – a workaround that scores well and reveals little. Strip that away and you're left with systems that struggle to solve a simple crossword puzzle, identify what's the same or different across two images, or navigate a basic 3D view. These tasks are essential to achieve human-level reasoning capability. Andrew will dig into why the current eval landscape systematically overstates capability, the structural reasons it does so, and how we got here from the viewpoint of someone who was inside a leading frontier lab. He will close with what he believes a more rigorous, consensus-driven eval framework needs to look like, and why the field needs to build one before the next generation of visual systems ships into the real world. Fixing visual reasoning starts with fixing how we measure it. For engineers building on top of these models today, whether that's document understanding, robotic perception, medical imaging, or any system where visual perception context matters, the cost of getting this wrong is already showing up in production.