I’m using “AI” here in the older sense, rather than the newer sense where it’s just another way to say LLM.

In this older sense, I don’t have anything against AI (even though I generally try to avoid LLMs). So, I thought I’d talk a little about the things I actually object to, when it comes to what people call AI these days. Specifically, what I object to (in order of objectionableness) are:

  • Using them to generate anything that looks like creative output. (It isn’t creative output, but it resembles it enough that I can waste a lot of time realizing that. That’s what I object to.)
  • The copyright theft at the base of LLMs. (I think half the profits (perhaps 40% of the gross revenues) of every AI company should be distributed to holders of the copyrights that were violated in the generation of the models).
  • The resource usage needed to run the inference engines. (Also the resource usage that went into doing the training, but that’s already sunk, so there’s no more point in complaining about it than there is in complaining about the resources that went into building your house.)
  • The fact that AI is unnecessarily used to do stuff that used to be better without it (such as web search).

I do also have some good thoughts. Generally speaking, there’s all kinds of stuff that (I hope) is going to get a lot better. Here’s an almost random sampling of ideas I’ve had. This list is most definitely not comprehensive. It’s not even the most important stuff. It’s just a few things I have been thinking of, because they’re things I want.

I would like an AI to keep track of everything I read (including whether I finish reading it, or give up part way through), and then (insted of trying to sell me something), guess what I’d like to read next. I’d pay money for this. (Not much money, but a little.)

I’d like an AI that picked up domain information what what I read. When I read an economics or finance article, I’d like it to put a little note over on the edge of the screen that I could click on, and then it would apply the information in the article to my situation. “That article, and three others that you’ve read in the past two weeks, suggest that European stocks might do better than U.S. stocks over the next year. Your portfolio is 43% U.S. stocks and only 16% European stocks. Click here for steps you could take to boost your European stock holdings.”

Of course, it should also track future results of each of those hypotheticals and compare them to both what I had before and what I actually did.

I’d like an AI to look at a blog post I’ve written and then from the taxonomy of categories and tags I’ve already created, suggest which ones I should use for that post. (There have long been “tag recommending” plugins for WordPress, but the last time I checked, none of them preferred the tags I’ve already got. Most of them seem intended for a completely different purpose from supporting your own internal tagging system. It seemed like maybe they were intended for finding keywords for maximizing ad revenue?)

I couldn’t think of a good picture for this post, but didn’t want to post it without a picture, so I thought I’d use this picture of my dog. It’s been hot here.

A dog panting, sprawled out on the carpet

There’s a lot of talk these days about the risks of AI, with many suggestions that it should be “regulated,” but with little specificity of what regulations would be appropriate. As usual, anybody who has an AI loves the idea of some sort of regulation, which would serve as a barrier to entry for competitors.

I have a suggestion that avoids that trap, minimizes the harm of regulation, and yet sharply constrains the opportunities for AIs to do bad stuff. It’s also easy to implement, because it requires little or no new legislation.

It’s very easy: enforce copyright laws.

Any firm that uses or makes available a large language model AI should be required to identify every copyrighted text used in training the model, and then share with all the copyright holders any revenues that the use or availability of the AI brings in.

This burdens existing AIs whose creators thought they were getting all their content for free by scraping the web for it, while giving a big leg up to any AIs that are simply trained on a corpus of text that the AI owner has the rights to. (I read about a physician who had been answering patient questions by email for twenty years training an AI on his numerous messages. He’d be fine.) That seems all to the good.

As to how much to pay the copyright holders, I think the publishing model of the past couple hundred years provides a good guide. Roughly speaking, book publishing contracts proved half the profits to the writer—but because it’s too easy to game the expenses side of the business to make the profits disappear, the contracts are written to provide something more like 10% to 15% of the gross revenues. That would probably be a reasonable place to start.

The huge cost of actually identifying each copyrighted text used, and finding the copyright owner is very much part of the desired outcome here: We don’t want people pointing at that difficulty and then saying, “Well obviously we should be able to just steal their work because it’s too much trouble to figure out who they are and divvy up the relatively small amount of money they’re due.” Making firms go through the process would provide a salutary lesson for others tempted to steal copyrighted material.