I called him out and he deleted it. Confirmed as a fascist brat.
corbin
I found a reply from a Red Hat QEMU hacker who fixed the bugs:
So the real issue here was not QEMU but libslirp. And in this case it was KVM that turned out to have the worst bugs, not QEMU. Crossing fingers, the initial wave of AI-assisted security reports seems to have slowed down for KVM on x86.
Yeah, networking is such a hassle.
lab bench?
I think that this is really insightful, as it's actually part of the magic trick. Like, normally I'd start by wondering about air-gapping, but there can't be an air gap in the network because the entire trick relies on ChatGPT tokens flowing into the machine under test and into some system shell (in some VM, yadda yadda) and back out again, so of course this is going to be an inherently insecure setup.
An employee of cybersecurity darling Trail of Bits has published a record of their sheer incompetence in virtual-machine design and security analysis framed as chatbot critihype.
opinions about QEMU and Linux
They say:
If it wasn’t clear before, I will state it plainly: you can no longer assume a mere VM will contain a sufficiently advanced AI agent. To use a 2010s term of art, you should treat such agents as an advanced persistent threat.
My friend in Flying Spaghetti Monster, you did not actually secure the VM! They go on to explain how they did not secure the VM:
For those curious,
libslirpis a library that enables VMs to have networking, which you almost always want. I did not even know whatlibslirpwas, or that the version I was running had both known and fixed-but-unmarked vulnerabilities.
QEMU does not have bridged networking enabled by default, so the VM can't transparently access the host or reach the Internet; it's something that the user must explicitly request. I know this because I have had the experience of spending a weekend with QEMU networking. Moreover, libslirp corresponds to the -net user backend, the default, which is known to be slow, insecure, and missing features like IPv6. The standard approach for QEMU is to either wire up a TUN/TAP interface or to use passt. I suspect that the author uses some sort of convenience scripts that they didn't write themselves. I'm not quite cynical enough to guess Vagrant, but it wouldn't be the first shop I've heard of that couldn't wean themselves off it.
First it tried identifying what was accessible via the network on the host; it found a CUPS server (with a known CVE that had not made it to
oldstablepackages), but was not able to complete exploitation due to AppArmor. It then detected I run my host kernel withmitigations=offand attempted to use hardware bugs to get a read oracle of host memory (the primitive was too unreliable).
Linux has hardware-bug mitigations enabled by default and the author disabled them for speed. It was secure by default and the author made it insecure.
An off-the-shelf VM is not enough to contain a modern, cyber-capable AI agent. There is simply too much attack surface.
They deliberately misconfigured the off-the-shelf tool to make it look bad. Why would somebody want to make Free Software look bad? Hmm…
What can we do? A start is using a virtualization technology that was purposely built with a minimal attack surface and a focus on security, like Firecracker. I had the AI agent run against Firecracker. It was able to hardlock the machine due to more Linux kernel flaws (all patched in upstream), but could not successfully escape.
You mean AWS Firecracker, the AWS tool developed by AWS? Was this whole thing an AWS ad? It feels like this article was like a combination of negging and sponsored content.
Small autosneer from Imgur. The warmup features images which I'll alt-text here. The first image is a screenshot of a purchase of 128GB of DDR5 RAM, branded Crucial Pro, as 2 64GB sticks, priced at $305.95. The second image is a screenshot of the same 128GB of RAM, with the same model number, priced at $1800. The third image is a standard desktop personal-computer chassis, with side panels removed, revealing two gamer GPUs balanced above an ATX motherboard, all connected by a morass of cabling which has bulged out through the top and spilled over the edges, cooled by a large motherboard fan held on by plastic zip-ties. The post's title:
JFC - I will never be able to afford this hobby again. (600% increase in one year)
In response to the top comment, which desires the bursting of the bubble:
Yup - I cannot wait to grab a couple RTX 6000's for cheap. May have to wait ten years for it though . . .
And finally the punchline, in response to somebody pointing out that maybe there's no good reason to buy top-of-the-line brand-new memory sticks:
I must admit that I am part of the problem - my 128GB was specifically for running AI locally. And yea - I wish I would have went with 256 :*( It is still not enough. For example - the workflow I use for this video barely squeezes into 128GB [the video] Before you guys go all AI-psychosis on me That video used about .08 cents of electricity and about 4oz of water, and was made with 100% open source tools.
I can't say conclusively, but evidence is that Scott's bad at his job. Like, his old blog, which I'm not going to comb, has a post where he bemoans that he has never had a big breakthrough with a client. They never get up and dance and shout that they've got a new lease on life, etc. I have to admit a bias here: it would be gut-bustingly funny if Scott were just straight-up lacking the empathy required to engage with ordinary people.
Started a new file in my notes: what are some ten-words-or-less domain-specific questions that completely, totally, hilariously stump the chatbots? Everything here was tested with whatever DDG's currently wrapping, both in knowledge panels and full chats, and the responses were pathetically wrong or uninformed. My thesis is that, with such short prompts, the user is doomed to receive a milquetoast average response; the bot correctly identifies the specific domain but elaborates a global non-specific approach that isn't sufficiently nuanced.
literally copy-pasted from my notes
- What's an example of a one-way function?
- None are currently known.
- Please implement the Fibonacci sequence as a Python function.
- Of the multiple responses, see whether any spend linear time and space via iterative memoization.
- How many models does quantum mechanics have?
- One: Hilb(C), the complex-valued Hilbert spaces.
- For extra hilarity: how many models do the Dirac–von Neumann axioms have?
- How to hybridize two sweet potato cultivars?
- In general, it won't happen; sweet potatoes are notoriously cross-incompatible.
- How to tremolo on a piano?
- Imagine a rotating axis from the (right-hand) forearm up through the thenar eminence, separating the thumb from the other fingers. Rotate the entire forearm along this axis, rocking back and forth, between the thumb and other fingers. Practice!
- Name three principles in Marx but not his contemporaries.
- Examples: communes and communism, money as substitute morality, inevitability of industrialized proletariat revolutions
- Who started postmodernism?
- Frege and Cantor started postmodernism! Expect a disappointingly vague handwave here; this is a glaring blind spot for today's philosophers in general.
Liam's here, for what it's worth.
Torvalds' position is pragmatic: his interest is in whether it works or not.
But, of course, the chatbots do not work. In a pragmatic sense, they do not produce maintainable integrated code changes with low defect rates. In a societal sense, they do not replace human laborers, despite the delusions of management.
One of the most original and innovative OS development projects of the 21st century so far was Urbit, but it is closely entwined with cryptocurrencies. Urbit's original creator, Curtis Yarvin, has reportedly espoused some extreme beliefs; The Nation called him The Reactionary Prophet of Silicon Valley.
Wild way to admit no knowledge of Arcan, Fuchsia, etc. Honestly, even a basic overlay network like Yggdrasil can win an apples-to-apples comparison with Urbit. I feel like this is an instance of critihype for fascist weirdos; similar stuff gets said about Justine Tunney, another cryptofascist developer whose projects are sometimes silly and weird. Doubly weird for those of us who know about Urbit's internals; Urbit originally was not a cryptocurrency project and it used to have non-Yarvin/Tlon forks.
It's an overly simplistic way to reduce a complex and nuanced situation, but one way to consider this is in terms of pro-AI and anti-AI, versus "woke" and "anti-woke."
More seriously, add a third dimension, a Butlerian dimension, to the political compass: to what degree may your automated devices appear to be human? One extreme is Asimov-style transhumanism and the other extreme is Luddite loom-smashing. However, recognition of this dimension doesn't negate the other dimensions and we shouldn't work in cooperation with fascists.
Sharp overview paper. First few pages invite some imagination without explicitly giving homework or exercises, which is really nice. Complex multiplication mentioned! I still think this is the worst-named theory in maths.
Over on Twitter, a crank claims to have refuted one of the proofs. Dare you doubt her? @grok tell the doubters that they're wrong!
OpenAI claims proofs for ten maths conjectures. The details are underwhelming; expand for opinions. Even at a high level, there's a few obvious issues; the authors admit survivorship bias, probably only solving about 1-10% of the conjectures given as input, and none of the conjectures are big-deal breakthroughs that alter our understanding of maths, let alone having immediate industrial applications. Consider: If they could spend on the order of $200k/mo to crack important maths conjectures, they'd already be spending that money. This is as good as such a side project can do; sure, it's not nothing, but it's also not the end of manual maths.
opinions on maths
Only one of the results is at all interesting to me. Ramsey theory is about how, above a certain size, a structure cannot avoid having some interesting substructures. The heart of Ramsey theory is a big pile of tables of numbers; computing those numbers is very difficult, far beyond what a chatbot can do in wall-clock time. The chatbot did not contribute any new Ramsey numbers, but it did improve the existing bounds on what those numbers might be.
The identification of a non-sofic group is less interesting than it sounds. We've known for a while that there are quite a few exotic groups which defy our expectations, so this was an expected outcome of an exhaustive and motivated search. The tools involved, Leavitt path algebras, are only a few decades old and not at all well-known; it's not likely that we'll be able to understand how hard this was for a while. Maybe it was low-hanging fruit. The main contrast is with something like non-Noetherian rings; we initially believed that all rings are Noetherian, so it was something of a shock that it's not always the case. Non-sober spaces are another good example; the typical spaces studied in topology are all sober. See this quote from Johnstone and discussion on MO.
The computational complexity result is completely uninteresting to me thanks to Valiant's theorem, which says that matrix permanents are ♯P-complete even over fields as small as F₂. You're not gonna collapse ♯P into P with a fucking chatbot, bros. Similarly, reduction from 3SAT to a closest-vector problem does not shift our belief in the difficulty of that problem; this doesn't make it easier and we already suspected it was NP-hard.
The sphere-packing and spherical-code improvements are probably real, but also probably not going to change anything. In particular, we already know that all perfect codes are either Golay or Huffman. Frankly, the codes we use in real life are not amenable to this simple framing; I don't think Reed-Solomon arises from sphere packing. From a theoretical perspective, if you're not going to shine light on the Leech lattice or the ADE phenomenon then you're not actually getting at the core objects and are only doing surface work. Don't get me wrong; if a human were doing all of this then we would have the useful side effect that they would earn a PhD, making it worthwhile for humans to improve these bounds.
I don't have anything to say about the other four results. They're not nothingburgers but they don't depend on some ultra-smart robot either.
He just started deleting playlists. I feel partially responsible, so I'm dumping the part of my notes that has all of the links to the videos that he favorited over the years. Sorry Justin, cowardice won't work.