Hacker Newsnew | past | comments | ask | show | jobs | submit | sinuhe69's commentslogin

Because it’s high time to rethink the incentives and rewards structures to nurture the next generations of mathematicians when corporate money and automated AI systems are mining the pool of limited interesting problems without giving back or enhance our understanding.

The problem they care about is not just “AI replaces human mathematicians” but rather the society thinks mathematicians can be replaced by AI.


As I understand you referred to Elon Musk. But this is Poh Shen Loh talking. If you don’t know who he is, look up or at least read his articles.

From his Wikipedia page:

> Po-Shen Loh (Chinese: 罗博深; pinyin: Luó Bóshēn; born June 18, 1982) is an American mathematician specializing in combinatorics. Loh teaches at Carnegie Mellon University, and from 2014 to 2023 served as the national coach of the United States' International Mathematical Olympiad team. He is the founder of educational websites Expii and Live, and lead developer of contact-tracing app NOVID.


weird bias by authority aside, do you really think that the daily decisions in this world today, and anytime soon, will be made based on the "let humanity flourish axiom"?

In daily decisions perhaps not, and certainly not when we were in the cave. But from time to time, we will face fundamental questions that need addressing clearly and decisively, and I believe it’s the case with AI. The more advances we make, the more frequent such questions will arise. That is also one of the main reasons why we don’t see any alien civilization. It’s getting harder and harder to keep the entropy low as we gain more knowledges and cognitive power.

How's humanity tackling that climate change question then? Lots of talk. Even lots of action I would argue. And yet, not enough. Why do you think that's going to change with discussions around AI when there are trillions of dollars at stake?

Because it will never change if you not try.

> Because it will never change if you not try.

That's why it's important to convince people to not try. It reduces the medium term risk to investor returns.


Irregular again! The single company that was responsible for the sloppy configurations and the hacks by OpenAI, Anthropic and now Google. This company and their partners should be held accountable for the crimes.

Yeah, something is amiss at that company. They've been at the center of all these.

If you get caught once and don’t stop. And then are caught a few more times, well… seems obvious these actions are taken with intent.

You use the low cost Mimo-V2.5 and not its big brother Mimo-V2.5 Pro? I also made good experience with Mimi-V2.5 when used in conjunction with prewalk mode. But then other models got so cheap and perform better, so I only use Mimo for background tasks.

The availability of Mimo over Openrouter got however, much worse recently.


If it works for you, great! Just keep going. I only have the nagging question that if it’s so easy, why would the clients not do it by themselves? Certainly they understand their problems better than you and they can adapt to rising issues faster if their in-house team does it themselves. And certainly what kind of AI you can access, they can too and perhaps even more. So what will remain your value propositions?

I’m a critical user, not a categorical denier. But there are certain categories where I see the current AI face a wall: fitting new ideas, or existing code into an existing architecture particularly. Or work on the backend when constraints are not strictly enforced (for various historical reasons). Schematic understanding is still an issue.

Sometimes additional chunkings and detailed planning/instructions suffice. But other time, humans are still the best vehicles to code.


> nagging question that if it’s so easy, why would the clients not do it by themselves?

It's easy if you're already an expert at software engineering and know how to leverage AI. For people like that, it's a phenomenal upgrade (this is true for me and several co-workers I chat with, all of whom are chasing cool ideas on side projects). But if you're non-technical, or not use to thinking about requirements, or think that using AI is "give me the prompt", it's a pretty big moat to cross.


> I only have the nagging question that if it’s so easy, why would the clients not do it by themselves?

Plenty of reasons, including:

1. They don't (yet) know how to use AI to accomplish what they need.

2. The ROI is still meaningfully positive and they don't want to have to do it themselves.

3. Having a third-party do the work provides protection for decision-makers. If the project fails, the third-party takes the blame and "nobody gets fired for buying IBM".

4. The vendor does bring valuable insight to the table and pairs it with the use of AI to deliver a result that wouldn't have been possible in-house.

This doesn't mean that everyone will be able to get work and maintain their rates in AI world but some people will.


I recently tried to encourage my smart, technical, but not-programmer friend to build an app using LLMs and he was just like, "I still have no idea where to even start."

People talk a lot about how "real programmers" used to have to clean up or actually deploy everyone's half-baked MS Access app, and that sort of thing will probably still be the case for a while.


> to clean up or actually deploy everyone's half-baked MS Access app

So we are switching to LLMs to be fucking miserable in our jobs?


The whole point of the MS Access reference is that similar situations have been tropes since at least the 1990s. Bad code generated by someone who doesn't know how to program well -- whether that person is supposed to be a professional programmer but is incompetent, or has a different job -- is nothing new, and neither is having competent programmers clean it up. LLMs probably generate more of it, but can also fix a lot of it, or at least patch it up.

A year ago, LLMs were not useful for me as a programmer. Now they are: the models are better, they can use long contexts more effectively, and the harnesses are better at helping the models. Nowadays my job is mostly not programming, but LLMs let me organize and prepare tools in spare time rather than needing days or weeks of attention. I would not trust them on a 200k+ line project -- and Claude Opus 5 has issues even on 50k LOC projects -- but they absolutely can help given good direction and a narrow enough scope.


Then I don't understand at all what point you were trying to make about being miserable in our jobs. Cleaning up after bad code has always been part of the job for decades; so has balancing the creation of new technical debt against resource availability.

> The whole point of the MS Access reference is that similar situations have been tropes since at least the 1990s.

I do actually understand that.

I have done this very put the Access database on the web job myself. (FWIW I was well-paid for it and the firm I worked for earned a fortune, but this was in 1997)


I mean they've boasted[1] about their generated code being so good that they add basically zero value and that they're riding out the rentseeking for as long as possible.

Meanwhile they've apparently played with LLM porting of Postgres (or Postgres features?) and "would not put it in production"[2], so we can also guess the kind and scope of features being discussed here.

[1] https://news.ycombinator.com/item?id=49594498

[2] https://news.ycombinator.com/item?id=48855951


They're not "adding no value", the SLA is the value. I have a business problem, and I pay you to solve it and support your solution for me, is a completely different proposition than standing up your own solution as a non-technical org and maintaining it.

> They're not "adding no value", the SLA is the value

The code their Astra+Fable setup, automatically generated from requests customers leave in voicemails, is apparently so good they never correct it. If the stories being woven are true, the SLA is literally just a middleman's 100k cut.

At the very least there might be some particulars here that aren't universally applicable.


>"the SLA is literally just a middleman's 100k cut"

Indeed, welcome to enterprise software.


The second was just for fun; it is fun. This is very far from the code we roll out at customers though. We build LoB apps which are basically crud apps (they are not but most here would call it that).

We can have fun and make money right? Possibly at the same time but not in that case.


> The second was just for fun; it is fun. This is very far from the code we roll out at customers though.

Which is fine, great even (I do the same thing!), but in trying different things outside your day job, it seems like you do understand why people working on different things than you might be "reporting all these negative AI experiences"?


It's a pretty bad sign when you need to resort to this kind of comment stalking to try to find ammunition for a general argument. This user could be a complete con man, and it still wouldn't invalidate the core thesis of LLMs providing business value.

> you need to resort to this kind of comment stalking to try to find ammunition for a general argument

It's not a general argument, they made specific claims, but vague posted (and implying everyone else must be crazy) to the point that the conversation is derailed by a bunch of people trying to figure out what they meant.

> it still wouldn't invalidate the core thesis of LLMs providing business value.

lol, no, "providing business value" is not what was claimed:

>> What are people here doing exactly that they are not riding the gravy train and even reporting all these negative AI experiences?


Hard disagree. It adds context to what they said.

Yes but it's cheaper?

Of course they will do it by themselves.

Because it's cheaper.

I mean it's weird that we all imagine reasons why we're still relevant when we have set fire to the thing that made us indispensable.

It's cheaper.


> I mean it's weird that we all imagine reasons why we're still relevant when we have set fire to the thing that made us indispensable.

Very few things in this world are all-or-nothing, and not every purchasing decision is based on price alone.

Do you always buy the cheapest meal? Car? When you renovate your house, do you always choose the cheapest contractor?

Tons of developers will lose their jobs, and many more will find it hard to maintain the salaries/rates the industry has been accustomed to. This is already happening. The days where an average graduate from a run-of-the-mill CompSci program or even a coding bootcamp could sleepwalk into a $150,000+/year entry-level job are largely gone. The days where you have job security simply because you're a competent developer with 10 years of experience are in the process of going away.

This does not mean that there is no subset of developers who cannot be successful in this market. There are people who are doing just fine because they know how to articulate their value and sell themselves to employers or clients.


> If it works for you, great! Just keep going. I only have the nagging question that if it’s so easy, why would the clients not do it by themselves?

Because it's not (yet) that easy, especially if we're talking a complex and genuinely useful app. Agentic coding is fast, but it's not magic.

> I’m a critical user, not a categorical denier. But there are certain categories where I see the current AI face a wall: fitting new ideas, or existing code into an existing architecture particularly.

Exactly. Which is why Joe Average still will have little to no luck vibe coding anything serious or novel. You still (IME) need a lot of active guidance, still need to push away from dead ends and propose alternative algorithms, and it takes hundreds of prompts to go from concept to what I would consider beta. (But, this is just my own experience, and it's possible I'm doing it wrong?)


I don't think that really matters. I have already seen companies generating their own code and just QA testing it, with no software engineers in sight. Or they ask us to "review" their 10k loc PR in 2 hours.

The spirit of the parent comment is true. People feel like it's productive and even if they make something worse, they are going to use it.


Clients are starting to, that’s exactly why I am not sure why AI doom deniers here think this will think it will go well for them. But it will still take awhile and we are enjoying that time. If we see new markets, we will move into them.

> I only have the nagging question that if it’s so easy, why would the clients not do it by themselves?

They can do it. What's the issue?


> I only have the nagging question that if it’s so easy, why would the clients not do it by themselves?

I mean, that's exactly what's starting to happen, we have more and more clients to whom we propose a quote and their answer is "guess i'll just vibe code it" or come to use with an app that they vibe-coded and does the job, and they're content with it. So far it seems to work out just fine for them.


Absolutely! I’ve never lied to anybody and any work we copied was purely accidental.


Snap is designed to be more expressive and powerful than Scratch. But I find debugging them is very painful. Changing the name of a variable or block for example, could create “holes” in the calling sites but the system can or cannot report an error and fails silently. Snap is IMO more flaky and the team seems more eager to add features than polish the existing ones or make the system mor robust and stable.


> Changing the name of a variable or block for example, could create “holes” in the calling sites

WTF. Isn't the whole point of visual programming that you are NOT bound by the limitations of text as a medium? A block refers to a specific variable. It shouldn't matter what it's called, it shouldn't matter if what it's called changes, it should still refer to the same variable even after renames.


That one is a bit more defensible for variables, when you rename variables, it gives you "rename" or "rename all". "rename" renames the variable, leaving uses of it the same, "rename all" also renames where you used it. Deleting a variable leaves the blocks that references it in place but they'll throw an error if called (IIRC this is useful if you wanted to switch from a global to local variable or vice versa).

Now, custom blocks have a different issue, where deleting one just severs any code that was using it (so if you had some custom blocks in a program: [on start] -> a -> b -> c, deleting the definition for b will mean c no longer gets called). There's now a block to delete custom blocks, so you can make self-destructing code.


They've fixed surprisingly many of the quirks, but my biggest gripe with Snap is that they don't like to add documentation with new features. A lot of the older content has a nice "help..." context menu option with a description of what it does (which is often out of date itself, e.g. the "split" and "join" by blocks feature isn't shown on those blocks), and newer blocks have the context menu option... but nothing pops up.


The documentation problem is entirely my fault. I managed to stay up to date through version 8.0, but then so many things happened so quickly that I got swamped. Then I got depressed and gave up.

Also I made the mistake of writing the manual in MS Word. At the time, I couldn't find a standard way to insert pictures into a TeX document, or I would have used that. (Now, of course, there is a standard way.) There is an effort underway to convert the manual into a web-based format in a git repo that anyone can contribute to, but that effort is 90% done and you know what that means! :)


Six months ago I was teaching a beginning-programming course to six kids, using Snap!, so I was reading the forums. And there was a group effort to add documentation to many of the blocks that were lacking it. I haven't been back in a few months, so I don't know how that effort is progressing, but there were definitely some undocumented blocks that were getting documentation completed as I watched, and eventually added to the Git repo to land in the next release.

So it's getting better.

The other thing to remember is that the core Snap! development team is just two guys, Brian Harvey and Jens Mönig. When they're focused on things like trying to figure out how to get macros into Snap!, so that Snap! can truly be a Lisp (right now it's only most of a Lisp), they tend to leave the documentation effort to the community. If they had a larger team I might fault them for that, but with just two guys, I can't really blame them for focusing their efforts on things the community is less able to do, and leaving things the community can do up to the community.


I have to clarify that the Snap! implementation is 99% the work of Jens Mönig. It was 100% Jens for the first two major releases, back when it was called BYOB ("build your own blocks") and was an extension of the Scratch source code. Then we wanted to use it in a CS course for non-majors at Berkeley so I got in touch with Jens and started complaining about missing features! The result was an intense collaboration in which the coding was still 100% Jens but I contributed ideas about user interface and features. I think my biggest contribution was teaching Jens about lambda!

Our team has officially grown to six people, adding Bernat Romagosa, Jadga Hügle, Michael Ball, and Joan i Pelegay. And several Snap! users have made major contributions, especially to libraries that extend the reach of the language. But the interpreter itself is still all Jens.


Glad to know the team is growing, it makes me slightly less worried about the project's bus factor.

"OpenRouter runs per-provider benchmarks on the same model: GPQA Diamond and TAU-Bench Airline (a tool-calling task)." why I haven't never seen it? Click on the link brings nothing! Benchmark is only for the model. Per provider is a performance matrix (latency, throughput) What did I miss?

--

Update: oh, that is the AutoExacto Benchmarks! Now I see it.


Then I guess you will need millions tones of perfect mass to light converter like Astrophage just to get to the neighbor. The universe is simply too big to comprehend.


Your benchmark is a curious one. I didn't see you included Muse Spark 1.3 contributor even though its price is much lower even than DeepSeek. The low price changes many recommendations completely. And FWIW, DeepSeek retain and train on your data, too.


I waited until NovitaAI provider became available on OpenRouter.

I have their guardrails enabled to not allow requests to providers that train on data.

As far as they say though...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: