Stop obsessing about the goddamn robot
Around the start of the year, before I had looked at any LLM seriously, I was talking to someone at a developer conference. He was despondent, because (in his telling) he had just got access to the newest Opus model and had simply prompted it “find and fix all my bugs”, and it did, and he was now wondering if he’d have a job before the year was out. That felt like a huge leap to make, but I assumed there was more to it and asked something like “interesting, I hope not, what makes you think that?”. In a moment he flipped, gave me a smug shit-eating smirk and said “oh, you’ll see”, and the harder I pushed for more information, the harder he defended the apparently growing LLM abilities and his prediction of doom for our industry. I actually wasn’t even disagreeing, I was more interested in his story. I dunno, maybe he just writes a lot of bugs?
More recently, I’ve been working on a review and redesign of OpenZFS’ admin permission and delegation system. A lot of this is searching multiple codebases, reading papers and books, developing use cases and personas and scenarios, and attacking the design looking for weaknesses. It’s interesting work, and very slow, and requires a lot of attention to detail. I told someone that I knew would be interested “I’m working on this” and he said “oh, just you?”, clearly implying that maybe I meant “the AI is doing it”. It’s not the only time I’ve heard some version of this, and though it’s not a big deal each time, when it’s said enough times it starts to carry weight and make me think more about what is actually going on under there.
And so I have been thinking, for most of the year, and I’ve largely concluded that LLMs are not creating new problems so much as making it easier to see the ones we already had. If you didn’t see them before, it’s easy to focus on LLMs as the cause, and miss that these are just the same old human problems we’ve always had, and that’s where our focus should be.
Please, everyone, stop obsessing about the goddamn robot.
We’ve had new and scary tech before: the internet, videogames, television, telephones, cars, looms, the wheel, fire. Each time, we had enthusiastic early adopters, doomsayers, and everyone in between1. Ultimately, we find a workable balance and most of the concerns fade into the background most of the time, and we manage the extremes.
I suggest that at least part of what makes a technology “good” or “bad” is the degree to which we focus on the technology itself vs the humans around it, and that more focus on the tech leads to bad outcomes, while focusing more on the humans leads to good outcomes. I suggest that cars and the history of transport are a useful analogy to LLMs.
Cars (and trucks and other related machines) have made long-distance on-demand travel possible, cheap and accessible for individuals in a way that wasn’t possible before2. They’ve allowed a form of independence, and with other technological and societal changes, they’ve become available to even the “lowest” members of society. They’ve made it possible for people to have more choice in where and when they work, live and play. They’ve enabled an incredibly high standard of life for much of the world, not least because everything from food and medicines to luxury items can easily be transported long distances from their place of manufacture. For some, like my disabled children, cars have opened up the world to them, both in increasing their physical access to the world, but also in creating a safe, isolated space in the world.
Of course, there are loads of downsides. We paved a lot of the planet to make it easier for cars to travel there, destroying large swathes of land to do so and quite often cutting through existing physical communities or bypassing them entirely, impacting local economies and livelihoods. There are direct and indirect environmental impacts from car manufacture, use and disposal. They are high risk, killing a lot of people directly and scores more indirectly, and that’s before you factor in the nightmare of oil wars.
Still, as a society, we’ve managed to settle on the idea that the benefits of cars are worth it, and we try to design or regulate away the worst parts. When we do a good job of this, we get public policy that supports and educates on ways of minimising car trips, road safety initiatives like speed controls, urban centres designed for walking and bicycle use, and so on. We focus on what’s best for humans, and reduce cars to a tool that has its place, but is not the point. On the other hand, when we do a bad job of this, we get road and freeway overbuild, huge and loud vehicles, environments where walking is outright dangerous, and so on. That’s not the fault of cars, as such, but rather, what we chose to build for.
LLMs seem to me to have a similar shape, or perhaps could have a similar shape. On the good side, they give access to not just raw information but actual knowledge in its context, which opens up whole fields to people who would not otherwise have access to them. We can have an idea for a computer program, app or website and see it realised without any particular technical knowledge. We can describe complex sets of medical issues and start to understand what’s happening in our own bodies without spending exorbitant amounts of time and money bouncing from specialist to specialist, desperately looking for help. We can make sense of complex legal interactions, or research obscure histories.
And yes, if we focus too much on the tech and what it could do, then we miss the actual harms to humans - overbuilding datacentres with their accompanying environmental costs, permitting overzealous managers to “replace” their employees, uneven intellectual property rules, privacy violations and misinformation.
The answer, then, is to focus on the humans. Understand the tech, educate people, set appropriate regulations and cultural norms and reduce the tech to just another tool in our box, that has its place but is not the point. We wouldn’t obsess about the robot more than any other piece of furniture.
So, coming back to my own interests, I see a lot of open source projects producing AI policies, and that ones that aren’t are getting plenty of questions about it. OpenZFS is one of those; at time of writing we don’t have an AI policy, but we’ll likely be considering it soon, so I’ve been trying to work out what I would want one to say, if I want one at all.
Debian recently decided on their own AI policy via their General Resolution process. The outcome was a proposal called “Responsible Use of Generative AI”. In part, it reads:
Debian neither endorses nor prohibits the use of generative AI tools in the development, maintenance, or documentation of software, packaging, documentation, and other media published within the Debian Project. We recognize that such tools can substantially improve the productivity of contributors when used responsibly, allowing volunteers to spend more of their limited time on work that requires technical expertise, judgment, review, and collaboration.
The Debian Project nevertheless expects that all contributions submitted to Debian, regardless of how and with which tools they were produced, satisfy the same standards of quality, correctness, maintainability, and legal compliance. The use of a generative AI tool does not diminish the contributor’s responsibility for the work they submit. Contributors are expected to understand, review, test, and, where appropriate, modify AI-assisted output before incorporating it into Debian. Blindly accepting or uploading AI-generated material without appropriate human review is inconsistent with Debian’s established development practices. We encourage our contributors to disclose whether a contribution was made with AI assistance, but do not require them to do so.
I like this and I’m glad it was the result. It basically says that you should use AI tools as little or as much as you want, but the standards your work is held to and your responsibility for it is on you, just as it always was. This works as a policy because Debian already have extensive documentation that describes their purpose and values, in particular the Debian Social Contract, that all contributors are required to agree to. So their AI policy is essentially “same rules as always”. In a way, they don’t even need a specific AI policy, as this is just an extrapolation of existing policy (not that having an explainer is a bad thing).
This, I think, is why I don’t like the idea of having to actually hash out an AI policy and particularly the idea that it might be contentious. If you already have a mission statement, set values, or similar guidance, then your AI policy should be mostly obvious from it, and if you don’t, then you can’t actually develop one because you have no way to measure it.
One of the most important presentations I’ve ever seen is “Platform as a Reflection of Values” (Cantrill, Node Summit 2017). It describes the importance of understanding and documenting your values as a project, community or organisation, that is, the things that are non-negotiable even at immense costs. The context is what happened to a project and its community in the lead up to and aftermath of identifying a divergence in values between the project and its corporate steward. Along the way, it presents a handful of technical projects, communities and organisations and suggests what their values might be3, and where you see it in interactions with the community and just studying how the software is constructed. And for a project or community that doesn’t have it written down and needs to, the template list discussed is a great place to start. In fact, it’s great for individuals as well - go through it, and decide what are the non-negotiables for the work you do and the projects you contribute to (and add your own, of course - this is not meant to be prescriptive or exhaustive).
OpenZFS does not have anything like this written down, but I don’t think it’s hard to at least make a start on what should be on a list like that. Things like “integrity” are easy; the whole point of the system is to make sure you can trust your data. Something like “security” is not on the list (as a non-negotiable, that is) but probably should be.
An interesting rule we do have is that all code contributions have a signoff, and that signoff must be by a real, actual human. We allow sponsorship tags in commit messages and corporate copyright statements, but we insist that our AUTHORS file only lists humans. Part of that is a protection against copyright & licensing rug-pulls; by making sure that everyone shares in the copyright, no one can ever unilaterally change the rules. The main purpose though is we consider that signoff to mean that a real human takes responsibility for the change and either owns the change outright or has permission to contribute it. That creates a constraint - if a human must vouch for a change, then a human must understand it (and especially if they want to sign off on the more complex technical values like “integrity”). And if a human understands it and could have written it by themselves, then I’m pretty ambivalent about what actually spat out the raw bytes.
This signoff rule also makes a statement, that this is a place for humans, which means we can start to admit human interactions into our thinking. OpenZFS has been around for 25 years, and I see no reason that it can’t exist for 25 more. This means it will outlive any individual, which then means we have to put maintainability and readability higher up on our lists of priorities. This presents further constraints for the use of tools that generate code - we need to be able to make changes that will be readable and understandable long after the author is gone.
I think it extends further than just the code. One important part of code review for me is the opportunity to teach and mentor another contributor, and in turn, listen and learn from them. This makes LLM-written PRs difficult, because they tend to be a vast information dump that doesn’t get to the point. The very worst part is when I reply, and I get more LLM in reply. This is bad, but it’s nothing to do with tools - this is entirely because I expected a human interaction there, and I don’t get one.
On the other side, we now get bug reports from people that don’t know anything much about the inner workings of their computer; they just got a crash on screen, asked an LLM to analyse it, and gave us the output. It’s hard for me to push back on this very far, because this is a real issue for someone, and they tried their best to report as much useful information as possible. People trying their best to work out what their own computer is doing and to make it as easy as possible for someone to help them is surely behaviour we want to encourage.
I don’t know how to write all of this into an AI policy that says “it’s ok for this, not ok for that”. I do know how to write it if I already have an overall value statement that says “don’t lose data”, “humans first” and “25 more years”. The AI “policy” would mostly be examples like the above, because the vibes we’re looking for are already spelled out.
I have been doing open source for a long, long time. The best version of it runs on mutual trust.
Tools can’t do trust. “My trusty hammer” is actually reliable, not “trusty”. When I pick it up, it’ll do the same thing it has done a thousand times before. The tool is incapable of intent towards me, good or bad. This is as true of an LLM as it is of any other software you use.
However, you can and do trust (or distrust) the humans that built the tool, wielded it, or stand behind its results. This has been true for as long as humans have been tool users, and doesn’t seem to be going anywhere.
This expectation of mutual trust is why the “oh, just you?” question has stuck with me longer than it probably should have. I work on software whose entire reason for existing is to keep your data - your life’s work, your memories, your livelihood - safe for the long term. I’ve spent years trying to earn the kind of trust that lets people rely on that. My first instinct to the question was to feel a moment of mild annoyance, then shrug and move on. But the more I heard it and the more I think about it, the more uncomfortable I feel about it. If the person saying it is really suggesting that a tool choice might be enough to bring all that earned trust into question, then either I haven’t actually built what I thought I’d built, or trust was never really what this relationship ran on in the first place.
So please, stop obsessing about the robot. Instead, focus on the humans in the room and how to make every interaction they have be the best one it can be.
-
I confess, I couldn’t find much about how people felt about fire at the time, but I can imagine the person in charge of keeping people warm wasn’t thrilled about it initially. ↩︎
-
Obviously cars were not the first generally-accessible personal transport option, but maybe the first of their kind. Horses have ongoing maintenance overheads that put them out of reach of most individuals; bicycles are cheap and accessible individual transport but only for short distances; trains can do long distance, but have less flexibility for choosing when and where to travel. ↩︎
-
ZFS is included in this list, but it should be taken as coincidence; it’s not the point of the talk and not why I recommend it. ↩︎