comments (10)

  • The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting.

    Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":

    https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f

    Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...

    simonw

  • I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:

    - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.

    - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.

    - Document parsing (extracting the relevant trip info from PDFs).

    If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.

    jampa

  • Currently top at https://deepswe.datacurve.ai - beating Opus 5!

    https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!

    Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

    mattlondon

  • Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents

    Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents

    (I think thinking level low is a regression on 3.8 compared to 3.7.)

    simonw

  • The most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic's flagships are still image-only.

    Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.

    simonw

  • Something maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability literally outperforms at least 2/3 Chinese students, no to mention those who speak Chinese. After all, the model speaks like a real humankind if you prompt it well. That's AGI guys

    EFLKumo

  • People have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good.

    These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).

    brap

  • Looks like the strategy of regular updates with incremental improvements is working out well. Interestingly, the biggest jump in Artificial Analysis Intelligence Index score is for reasoning level Medium ( 3.7 was 51, 53, 57 for Low, Medium and High, 3.8 is 52,57, 59 respectively). I think scores at lower reasoning levels are more indicative of model capability since higher reasoning levels are focussed on benchmaxxing. We use the lowest reasoning level in production with good results.

    a11r

  • Wow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC?

    I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker!

    At this point it is a meme of course, but where is 3.5 Pro :)

    mattlondon

  • "The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)."

    Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

    j-bu