comments (10)

  • I heard a radio spot recently and I wondered if the voice was a real person or AI. It makes we wonder how such industries are dealing with this gen-AI revolution. We spend a lot of time here thinking about how it affects software developers, but I hardly ever see any commentary on how it is affecting screen and voice actors.

    petcat

  • Prompt engineering tip for Google employees: just add "P.S. Make sure the page works in Firefox too."

    037

  • > What stands out about Gemini Omni Flash is its accuracy: the details hold up under scrutiny.

    Quote under the video of a short Argentinian footballer wearing no 10 with "RESSC" on his back. Can't make it up.

    petegleeson

  • Interesting that OpenAI abandoned Sora entirely but Google are continuing to invest heavily in their own video generation.

    Maybe because they see video generation as key to developing "world models"?

    simonw

  • Anyone else becoming numb to these updates?

    I feel like I should be excited about being able to generate almost perfect videos but, I just don't care anymore.

    bluerooibos

  • Google does anything except launch a new version of Gemini Pro.

    guilhermeasper

  • I let myself get mildly excited with the last Omni release, but it turns out it (and this one) can't do the one practical thing I want - Sync generated video to provided pre-existing audio.

    Meanwhile, I'm happily using Minimax H3 locally on my 12Gb 4070RTX to finally finish the lip syncing to recorded dialog on my abandoned 20 year old Flash animation hobby projects.

    Nihilartikel

  • Before AI, cool and interesting shots carried the promise that it was reality - even if it was perhaps exaggerated.

    The amazing videos and photos carried the promise that I could experience that for real. They were aspirational.

    Today I suspect every cool shot is made of pixels arranged on a 2D screen by an algorithm. It doesn't do it for me.

    That makes me sad...

    DataDive

  • It certainly makes for easy demos, but I always struggle with the practical application. As in, what work or enjoyment does someone actually get from this? Ads and media pre production seem plausible, but it fails the 'how can this enrich life' in a way most other AI tools don't. Maybe for them that's not a consideration, if their only interest is the other meaning of enrich that might flow from ads and numbing rivers of slop.

    Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.

    I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).

    rcr-anti

  • Draft videos more efficiently in 360p

    While it sounds great you're quickly disappointed after you run the same prompt at standard resolution only to get a different result because it's non deterministic.

    cube00