this post was submitted on 13 Sep 2026
140 points (94.9% liked)

Technology

88072 readers
3326 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] melfie@lemmy.zip 1 points 1 day ago (1 children)

Looking at LLM benchmarks over time, there was still doubling of benchmark scores as of 2024 to mid-2025, then it inflected to more modest, incremental improvements, so sigmoid seems about right. Now we are seeing open models rapidly catching up now that the frontier has significantly slowed.

[–] MangoCats@feddit.it 1 points 1 day ago (1 children)

Yeah, improving on 2024 performance was a pretty low bar to clear, by mid-2025 I saw the continuing improvement and decided that even if it was marginal at the time, learning how to use it was probably worthwhile given the improvements that seemed to be coming. Those improvements definitely did come from my perspective. Much of it was in the harnesses - many things I used to have to tell models explicitly, repeatedly in fall of 2025 they started doing without explicit prompting by spring of 2026. I also think I learned what they could be expected to do well and what was a waste of time trying which made me more productive with them as well.

I recently heard that Kimi v3 has made significant progress in code quality - with some people calling it "on par" with Claude. I primarily use Claude Opus - when I tried Fable during their free preview it "felt" even better, but not enough better to shell out a lot of extra cash for personal playtime projects. Kimi isn't as accessible under the fixed monthly price model so I'm unlikely to try it anytime soon.

[–] melfie@lemmy.zip 1 points 1 day ago (1 children)

I saw this article today summarizing a Mozilla report that open weight models are about 4 months behind the frontier: https://arstechnica.com/ai/2026/09/exclusive-open-chinese-models-close-gap-with-silicon-valleys-frontier-ai-models/

Progress has slowed at the frontier and open weight is right on their heels. The general advice is that Fable and the like are best for niche use cases and to make an open weight model your daily driver because they’re almost as good and are a lot cheaper. In my case for personal use, I’m perfectly happy with Qwen 3.8 27B running locally and don’t care to deal with data collection or huge bills to get a bit stronger of a model.

[–] MangoCats@feddit.it 2 points 1 day ago

To switch from Claude to Kimi in service forms available to me would be a 5x price increase for how I use Claude (milking the subscription 7 day token limit dry after 5 days pretty consistently.) Decent self-hosting solutions seem to be running around $20K+ in capital equipment and more than $20 per month in electricity costs alone.