Generative AI: Excessive Reinforcement Learning

Do you know that without human intervention the Generative AI responses would be too toxic after training the model from the garbage on Internet? Do you know that OpenAI hired Kenyan workers to make ChatGPT less toxic? To throw in some jargon, this is called RLHF (reinforcement learning with human feedback).

Not sure who Microsoft hired to cleanse copilots, but apparently they went overboard. Here is a Windows Copilot jerk-knee response that I got to a prompt asking when to open with 2 No Trump in contract bridge.

I see, Generative AI refuse to generate since “trump” is a taboo topic nowadays. AI dummy!

Generative AI: Excessive Reinforcement Learning

Follow Us

Subscribe to our quarterly newsletter

Categories

You might also like

Follow Us

Subscribe to our quarterly newsletter

Categories