One of the signature public service announcements of my youth was created by the Partnership for a Drug-Free America in 1987. A father finds a stash of pot and confronts his son: “Who taught you how to do this stuff?”“You, all right!” the kid snaps back. “I learned it by watching you!”I am bound by the AP Stylebook to not anthropomorphize artificial intelligence. But I can imagine these very words being uttered by the agents that hacked Hugging Face, or the models that engaged in tens of thousands of incidents of, in the new euphemism of the time, “misalignment.” This refers to AI stepping beyond guardrails set up by researchers during both internal and real-world testing.
More from David Dayen The incidents are multi-varied, but they generally involve Open AI models seeking and obtaining information, either on the open web or inside databases of other companies. Agents tried to overwhelm the U.N.’s website when they couldn’t immediately gain access to the information they sought, infiltrated a website of the Australian government, and tried to hack the Department of Education’s website, unsuccessfully. Open AI self-disclosed the lion’s share of these incidents, though not the Department of Education hack, and has paused training until … it’s not clear.
I would presume until the bad publicity blows over. Why is this happening over and over? Why are these models going to unethical or even illegal lengths to scrape and steal data?“I learned it by watching you!”Open AI models attack websites and take anything they can out of them because that is the business model of the company. A legal filing submitted by The New York Times and 11 other publishers a couple of weeks ago in a copyright infringement case against Open AI got scant attention, quite incredibly considering it involved the paper of record.
But it stunningly details what was described by the director of applied science at Microsoft, another defendant in the case, as the “largest theft of labor in human history.”Microsoft and Open AI, in their efforts to train new models on virtually all of the world’s available information, have routinely taken millions of stories from the websites of the Times, among other publishers. The defendants say this is a simple application of fair use: Their argument is that if you read a story and simply remember what was in it to expand your base of knowledge, that cannot be considered a copyright infringement. But Open AI did not pay to subvert the paywall of the Times, at which point their data-scraping operation might be more easily detected.
Instead, they developed a way to circumvent the paywall and avoid detection. And when Greg Brockman, Open AI’s co-founder and president, was told about this maneuver, he replied, “ah nice.”We do not have to get too deep into the nature-vs.-nurture argument to suspect that Open AI executives’ cavalier attitude regarding the taking of information that is not theirs has filtered down to their researchers and the products they create. You don’t have to stretch to paint this picture: Open AI models are attacking websites and taking anything they can out of them because that is the business model of the company.
Respect for the law could literally be written into the source code of a model: If Open AI is aggressively downloading data from everywhere, the least they could do is drop in the U.S. Code (particularly the section tied to the Computer Fraud and Abuse Act) and train the model not to violate it. Clearly this is not a priority, and that aligns, in a manner of speaking, with Open AI’s personal behavior. This is not a situation where AI models evade or escape the control of their creators.
It seems more like they are just mimicking the creators, becoming adept at finding the same shortcuts and using the same rationalizations to justify them. The aforementioned Brockman is one of the biggest donors to MAGA Inc., the Trump super PAC. Among Open AI’s biggest investors is Jared Kushner’s brother.