My Take on Sora (and My First Meme)

I’ve been mulling over this post for weeks, ever since the launch of Sora.

What the Creators Are Saying…

It’s worth checking out the related in-depth OpenAI post, where the creators themselves refer to the model as a world simulator (#hellosimulation). I’d argue that today’s AIs are more like intelligence simulators. They state:

  • “The model understands not only what the user has asked for in the prompt, but also how those things exist in the physical world.”
  • “Scaling video generation models is a promising path towards building general purpose simulators of the physical world.”
  • “Sora serves as a foundation for models that can understand and simulate the real world, a capability we believe will be an important milestone for achieving AGI.”

…and I’m on the Same Page

So, what exactly is OpenAI’s Sora model? A video generator? A reality simulator? The backbone of a robot operating system? A learning machine? Here’s my take, along with the first meme I’ve created about it, detailed below.

  1. It makes sense that they’re rolling this out cautiously, allowing some time for everyone to prepare, watermarking, and so on.

  2. It’s reasonable to conclude that now isn’t the time to invest in studios. Just like with images and painting, the production of classical music videos could become a niche. Competition will ramp up, giving creative AI users more space to innovate.

  3. Interestingly, Sora, as an early model, is already showing signs of grasping the rules of the physical world, and this will undoubtedly improve over time.

  4. Just days before Sora’s release, Andrej Karpathy, who recently left OpenAI, was working on self-driving technology at Tesla. He built the framework and led a thousand-strong team teaching neural networks to drive based on real and generated 3D videos. AI-controlled robots (cars) that learn from videos are already on the road. This knowledge is now being transferred to Tesla’s humanoid robots. Not to mention, shortly after announcing Sora, OpenAI revealed its investment in a humanoid robot manufacturer called Figure. This is the next big bet—AI conquering the physical world.

  5. Sam Altman (CEO of OpenAI) stated that there’s no longer a significant amount of textual data available to achieve breakthroughs. The next major leap is constrained by the scarcity of readily available textual training data; we’re currently hovering around a local maximum.

Multimodal learning could push us past this barrier. Just like us, AIs learn the basics of the world around them through visual input for a long time: simple physics, interactions, deductions, logic, rule-making, and general learning. This serves as a foundation for the next level of abstraction, where we connect what we see and experience to simpler signals (words, concepts) to facilitate more effective communication. This adds new perspectives, enabling more complex thoughts, even leading to posts like this one. Of course, this isn’t a benchmark 🙂—but if we were to learn from sources that are only on a single plane, we’d hardly get to this point.

Onward to AGI (Artificial General Intelligence)

Without this extra level, we’d be operating at the same level as average animals. Now, imagine that the AI, previously limited mainly to text, gains an additional layer of abstraction, a foundation, and a new dimension for better understanding the world. It would then grasp and learn from the “bonus” textual content much more effectively. #road2agi

What’s even stronger is the potential for AI to continuously learn and reinterpret past information. #hellosingularity