Privacy
Icon base by Lorc under CC BY 3.0 with modifications to add a gradient
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
Wow I had suspected this would eventually happen, but for it to be this blatant this quickly is alarming. They feel so emboldened due to the complete lack of consequences for stealing.
Surprise surprise...
Further reading: "Is OpenAI Taking Everyone for Fools?"
https://read.misalignedmag.com/is-openai-taking-everyone-for-fools-2481fa851544
They also spend the equivalent of something like 15% of that university's mathematician research budget for the past 50 years to do this one proof, using plagiarized data to do so.
while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
So either they don't track what goes in their training data, or they do but their LLM is unable to precisely credit sources. Maybe a bit of both. That's convenient, this way they claim great discoveries without crediting all prior work they're relying on.
LLMs are a great plausible deniability generator. These allow one to plagiarize while saying with a straight face no one knows what source material the result are based on.
