Hacker Newsnew | past | comments | ask | show | jobs | submit | sigmar's commentslogin

>I’ve emailed some of the mathematicians with a few proposed typo fixes, and I got confirmation that at least a few of those fixes seemed real. However, some of the problems that weren’t backed by Lean also turned out to be misunderstandings.

I think this project is really neat, but is it appropriate to cold email specialists before you've put in enough hours of effort to describe yourself as more than an "amateur"? OP's emails may have been helpful, but billions of people use these LLMs to wade into new areas and email is already low signal-to-noise.


Yeah it's a pretty tough question! I've resisted doing that until I had a relatively high certainty that their published results contained minor mistakes, which I assumed they would want to know about. I've also been explicitly apologetic and tried to keep it super brief.

A counterpoint - I was told a story by one of my professors in the late 1990s, about one of his professors -- he'd written a thesis, gotten hired somewhere like Princeton, and taught there for a few years as Dr. <Somebody>. One day he received a letter pointing out a construction flaw in his thesis. He brought it to the department head who read the letter, and said "Well, Mr. Somebody, ..." Ultimately he fixed the proof.

Upshot, if there are real errors in published work, I think most mathematicians want to know about them.


Yea but that was a person who actually put in work and had to think about and understand the problem. They didn't just generate something with a magic box.

As a software engineer, I could not care less about how a bug was found or by whom, as long as I can quickly verify it is correct, I always appreciate being able to improve my work. I don't see why it should be different for mathematicians (the ones I knew would think similarly, I would assume.)

If I may offer an analogy, what I tried to do is essentially reporting a bug after having a failing unit test that exercises the public API, without looking into the black box of internals.

> They didn't just generate something with a magic box.

It doesn't matter who found the error nor how it was found. An error is an error.


I'm definitely more pro ad than the average HN user (I think they occasionally show me products I want to buy), but this seems bad and one step closer to "the LLM is biased by openai's financial interests."

In the example of a "sponsored agent" at the top, how many people are going to think "chat with us" is just a way to continue the conversation? How soon before openai removes the friction of clicking that entirely?

this is what openai said before: "Ads do not influence ChatGPT’s answers. Ads run on separate systems from our chat model, and advertisers have no ability to shape, rank, or alter ChatGPT’s responses. Ads are separate and clearly labeled." https://help.openai.com/en/articles/20001047-ads-in-chatgpt

Do they need to rewrite that section? Because in the video there's an LLM chatbot inside the chatgpt app writing answers on behalf of "Heirloom grocery"


Disagree. I pitched this exact model to Google Search when we first started making LLMs. The idea was the bots were separate clearly from the core model so you’ll know when you were having a branded experience or not. And it s useful beyond ads. It can be for customer service or MCP interaction with the service directly.

yeah, I can see it as useful in customer service, but they're adding these sponsored agents directly in the chatgpt app. the word "ad" or "sponsored" isn't visible after you click "chat with us" at 0:15 in the video, only "learn more about business chats" which is vague. and if you scroll down you won't even see that, if someone switches apps for 20 minutes are they going to remember that this chat window is an ad?

It’s not an ad. It’s a business agent. The entrance is an ad.

This isn’t that hard to understand.


>This is reportedly a sort of "Enron" accounting which excludes some really big expenses like revenue sharing, the cost of model training and hardware deploymments

source? this seems false. reportedly the adjusted profitability includes inference and amortized training costs


source?

Listed at the end of my post.

this seems false.

Source showing this in accordance with GAAP (Generally Acceptable Accounting Practices)?


that says only that the gross margin calculation excludes profit sharing and training. You should read it more carefully

edit: def not gaap profitable or they would have said that to investors. and their stock-based comp is surely astronomically high on paper.


>lists a few methods that are quite similar to what’s proposed in the comments under that post, except for those two comments.

Is today the first time you've heard of a book cipher? Those blog comments didn't provide much progress.


>Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.

Reasoning about how to write secure software uses the same knowledge as reasoning about how to break/hack it.


Just remove anything software-related from the training dataset.

Which also solves the alignment problem with those who do not enjoy seeing LLMs write software. But that brings us back to: Aligned to whom?


I don’t buy it.

Reasonable about building secure software can take the form “this memory access might be out of bounds — that MUST be fixed” or “this process has access to an inappropriate privilege — this is a serious weakness”.

Exploiting things and the capabilities that the labs call “cyber” are about the ability to (a) find the issues mentioned above and then (b) string issues together and avoid all the imperfect mitigations to actually compromise something. That latter part was IMO not actually necessary to train extensively, and I’d be quite happy to use a model that has no special skills in this regard but that would do (a) without complaining.


Lots of private benchmarks already exist, where you have to trust the tester (ex Artificial Analysis, Arc-agi).

>To understand the truth about the Hugging Face hack, you could do a lot worse than to listen to Ed Zitron and Cal Newport's recent podcast conversation

Lol, okay...

The crux of this piece is Doctorow saying that the hack was just a stochastic parrot repeating steps it has been trained on. Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened. "Oh, it only made those paperclips because it saw instructions on making paper clips in the training data." These are some 2024 arguments...


I can't help but notice your counter arguments are "lol, okay" and "These are some 2024 arguments".

Not exactly convincing stuff.


Sure, if you remove almost all of my comment my argument disappears.

To spell things out for people that don't know about the topic: Zitron is neither an expert on the topic, nor a credible source of information: https://techreport.ngo/ai-ml/how-accurate-have-ed-zitron-s-a...

The paperclip maximizer is a thought experiment. If I tell an AI to start producing paperclips, it might start producing paperclips by doing unintended things. Technically, recycling the metal from all the world's bridges would assist in making more paperclips, but I never intended that. That's where my analogy to the post comes from- Openai intended the model to hack, but did not intend for it to hack HF. Do you think the fact 'recycling metal into paperclips' may have been in the training data is relevant to the thought experiment? Similarly here, it's orthogonal to the lessons from the HF hack and only brought up here seemingly to make it a fight over whether the models are really "autonomous"


> Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened.

Seems like a pretty clear and succinct argument. If your only counter is to complain about tone, you’re losing.


That is all that Ed Zitron deserves lol. Completely unserious person. Really on the same intellectual level as the flat earthers at this point. And at both we may simply point and laugh.

It’s shared context for anyone familiar with the field.

Amusimgly, also how cults and fascism works.

They also breathe air. What’s your point? Don’t exercise reading comprehension because bad people do it too?

Yeah, basically any human group has some shared social context. Most cults and fascists probably eat together sometimes too, but that doesn't make it a red flag.

What is a better/more faithful articulation of the incident in your mind?


Ok; but the other side is equally delusional.

> "Oh we told the AI to use the tools, as well as to not use those tools. It chose to use the tools - we consider this cheating (for neabulous reasons), so lets get everybody in a panic about the morality and ethics, and how we can program those into the AI."

We know perfectly well how to constraint these programs. Attack isn't growing faster than defense. The people who believe in existential risk and want to teach AI's to be nice, as the last line of defense aren't helping at all. They're just jumping on the fearmongering bandwagon, perpetuating an "other consciousness" misunderstanding of the tool.

I've not seen LLMs display competence we should be fearful of the damage _it_ will do if left unchecked. All the damage will be done by ourselves to ourselves, regardless of the safeguards ideas being floated about.

My current belief is this whole HF media circus started with the simple human desire of OpenAI engineers to frame it such, that nobody would question their incompetence & liability & complicity.

Nobody is ever held responsible for out of control forces of natural powers after all.


> The people who believe in existential risk and want to teach AI's to be nice, as the last line of defense aren't helping at all. They're just jumping on the fearmongering bandwagon, perpetuating an "other consciousness" misunderstanding of the tool.

It's not clear what "other consciousness" should be, but I presume it's misalignment. This is a recognized phenomenon, not fake news (that is, it can exist, I'm not implying it necessarily develops/spreads).

No doubt that right now there's no existential risk, but assuming that AI will be enormously smarter in the future, it's a valid point to doubt whether we'll be able to prevent/recognize/contain misalignment or not (even if misalignment won't be spontaneous, bad actors will surely actively develop it).

> Ok; but the other side is equally delusional.

AI development pointing to superhuman cognition, and, on the other hand, possibility of misalignment, are real. Put together with the fact that AI will be ubiquitous in the future, and there the disastrous scenario becomes plausible.


In the podcast, Cal Newport repeats at least 20 times that agentic AI is nothing more than an LLM execution loop, as if this is some kind of trump card, taking for granted the usual stochastic parrot trope. At no point does he even consider the possibility that the LLM output itself is noteworthy. These people are permanently stuck in November 2022.

At this point, the "stochastic parrot" people are sounding like they need their prng seeded with less predictable numbers.

The output hasn't changed because the input hasn't changed. The shoe still fits.

So far everything an llm has ever done is still consistent with fitting bits of training data together.

It looks like people because it is replaying things done by people.

Similarly in the other direction, the fact that people can and often do mechanical things (make bad art, follow routines, etc) does not prove that people are no different than machines either.

Just because there are these two overlaps in both directions doesn't excuse getting them actually confused.

I don't think anyone lacking the perception to distinguish these things simply because there are overlaps and similar appearances is in a great position from which to be calling anyone else stochastic.


The llm-driven tools of today are completely different from what they were 4 years ago. It's like the naysayers can't keep up with the development and decided to freeze time and attack the situation from back when it was still possible to be cynical and not wrong.

The llm-driven tools of today are just a faster toaster. It's like the credulous can't recognize a toaster when they see one and don't understand that a faster toaster is still just a toaster.

>or there is a more deliberative approach to assigning credit than who was "first" to solve some problem

it feels like an unintended consequence of the millennium prize is that people view the [last contributor to the solution] as the only one to make progress on the problem. I've never viewed Poincaré as solved by one person and the objective of the prize was to encourage more people to make attempts and contribute towards progress.

this issue is independent, but in these circumstances perhaps interweaved, with the 'ai is taking over math' concerns


Were the agents ever tasked with algorithm improvements? Post just says he didn't find any ("report essentially no algorithm advancements"). These LLMs are useful for optimization tasks where they can attempt a change and then measure performance boosts, so just wondering.

>We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations

Do you think it is possible that better math will lead to better physics models?


It might but the math results from GenAI so far have been limited to finding counterexamples to known conjectures, not building new mathematics.

> results from GenAI so far have been limited to finding counterexamples

Not all.

Ehrhart’s volume conjecture

Quantum parallel repetition for general two-player quantum games

Erdős Problem #183 on multicolor Ramsey numbers

Erdős–Sárközy Problem #12(i)/(ii)

Erdős Problem #125

Log-concavity of codimension-3, type-2 pure O-sequences

Optimal O(1/t) last-iterate convergence for Anchored Gradient Descent-Ascent


True, these results are from last month. My info was a little out of date.

Prior updated :)


Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: