> "it's not just a next-token predictor because a bunch of the training isn't about predicting the next token."
clever procedures on top of the base transformer architecture.
i used simplified words/phrases to summarise the same thing you two were saying (the intent being: here's a version that may be digestible when discussing with others).
apparently that means i'm wrong though, no idea why because it seems you've decided to be dismissive rather than constructively elaborate on why this simplified and digestible version might be wrong :shrug:
But it's a policy learned from a next-token prediction task. You could also call it an inferrer or generator or whatever. The point is that it takes as input a sequence of preceding tokens and emits one more token to continue the sequence.
So? tokens are emitted one at a time according to the output distribution & sampler algorithm, and the next token distribution is a function of the preceding token sequence only. The process by which the output distribution is shaped doesn't change the core mental model, and doesn't reduce its value. It's a prediction in the jargonic sense that an inference about future values of a time series is broadly called "prediction", and it's relevant for reasoning about LLMs because they are fundamentally limited to converting tokens sequences into next-token predictive distributions, and that bears on how they can do what they do and what their limitations are. Nothing about the training process changes that.
prediction is a very specific term of art in the field of machine learning. generally speaking, machine learning models like LLMs are based on probability; performing a statistical prediction of the likely y given some input x
Probability(y | x)
that's why we refer to outputs as a prediction. it is likelihoods and stuff. the output is never definitely correct as we're not dealing with heuristic processes.
> Prediction implies there is some "truth" or event or something that you can test against
there absolutely is a ground truth during training. the core predict-the-next-most-likely-token part of an LLM has a ground truth next-token. that's why you don't end up with generated text like: fish spurious send cattle chocolate phone happy meaning ball orange board canada.
> optimizes to predict the next token in training data
that is the optimization goal in training the next-most-likely-token core of an LLM, it basically translates to maximise the likelihood of predicting the next token x_i given the previous tokens
Read through the article and comments. You are talking solely about pre-training. I'm talking about post training.
Respectfully, you are miles out of your depth. GPT-2 didn't use any reinforcement learning and is often given as a toy example. That release was 2019 and models now go through a various phases of training with different objective functions and optimizers.
From GP, i.e. the context for this local part of the thread
> Autoregressive LLMs generate tokens one at a time, disputing this is just plain wrong.
next-token prediction i.e. the bit built during pre-training.
at no point in your reply to GP did you specify that you were referring to post-training. respectfully, it seems like this one is on you pal :shrug:
> GPT-2 didn't use any reinforcement learning and is often given as a toy example. That release was 2019 and models now go through a various phases of training with different objective functions and optimizers.
yeah. so? the toy example works for pre-training. see above.
All modern LLMs that actually get used go through post-training. The finished product is something which has been through post training. So they are not next token prediction machines.
> The finished product is something which has been through post training.
again, the finished product wasn't what was discussed by GP, and you didn't clarify that you were switching to discussing RL (which is still probabilistic btw)
You are conflating "half built" with "a piece of a system".
The model weights change as the model goes through the training process. They aren't stored after pre-training is done and other weights are put somewhere else. It's more like pottery - the thing changes. It's not correct to say something is soft and malleable because it once was.
> The model weights change as the model goes through the training process.
Yes. They do. You are absolutely right about that.
But the model architecture doesn't change as a result of the training process. A piston doesn't suddenly turn into a digital watch as a result of tuning an engine. Similarly, the transformer part of a GPT model doesn't suddenly turn into something else as a result of optimizing a loss function.
Just skimming through here but I think you have the wrong ideas with llms, I’d recommend Andrew Ngs course (correct me if you’ve already seen it or something similar).
I am not an expert, but I do understand the distinction that is being made here. It makes sense to describe the result of pre-training as a ‘next token’ predictor as that’s what it’s been trained to do, not because it’s an autoregressive architecture that produces tokens one at a time.
If this base is then trained using RL towards a different objective (maths and coding), the model becomes fundamentally a different thing and the recent models are clear evidence of that, regardless of they fact they remain autoregressive.
So, this is the cause of the problem.... People take an intro to LLMs course, follow happily along, and don't realize there is more to it than the next token prediction. And those courses teach how LLMs were built in 2017-2020 maybe. Then RL got added to the mix. The current models really are very different to the models from then - everything that is now considered "post-training" isn't doing next token prediction.
When you are done with the section on RLVR, consider whether the model is predicting tokens, or making moves. There is a reason the word "policy" is used in RL.
You're using the fact the both parts of training affect the same weights to support your argument that they're making the system do something fundamentally different after RL?
> The agents continue to poke around on DSEWiki. A few hours after they find the site, they start probing it for cross-site scripting (XSS) vulnerabilities. [...] The agent swarm starts testing whether they can execute JavaScript that they embed into the search page, and continue to do this for a few days
either the agents were doing free security testing for the site and “forgot” to submit a report, or they were trying XSS to gain something they didn’t have permission/authorization for.
also
> Hijacking: To take control of (something) without permission or authorization and use it for one's own purposes.
a mod had to go through and mass delete a bunch of pages that didn't belong on the site. no-one from the wiki site gave the agents permission to use their site as a message board. hijacking isn't being used here in the sense of "gained admin privileges to run crypto scripts" -- there are multiple ways to use a word.
I had a great experience using $PROJECT_NAME. Another sentence that says nothing at all; but look, fancy and technically correct grammar use! You’re absolutely correct in agreeing with the author. Of course there’s an opposing point to consider — but that’s not it, this is.
1a. Illegal immigration has nothing to do with visa based, legal immigration. illegal immigration is the type of immigration people are upset about here in the UK. most people have no problem with legal migration.
1b. anyone claiming the UK’s legal migration is any sort of problem right now, or in the past, should probably not be listened to in any way shape or form for anything they have to say on the matter.
2. a lot of the folks who have illegally entered the country end up asking for asylum, and by law their claims must be processed
3. the dublin convention states something along the lines of “asylum requests must be processed in the first EU country the person entered into”, ie it’s possible to remove people claiming asylum if they initially arrived in a different EU member state country
4. most people get to the uk having traveled through EU member states, because geography.
5. we left the EU, we cannot do 3 anymore. so we have a bunch of people stuck in hotels waiting for asylum claims to be processed.
illegal immigration got worse because we left the eu.
People absolutely do have a problem with legal immigration, especially post-Boriswave.
If legal immigration means millions of new people every year and rapid demographic change, of course that will cause a political issue.
It would be very strange (basically unheard of in human history) if it didn't.
People I know who are non-white all report increased racism (mostly not aggressive but noticeable) since Covid, despite having English accents so clearly not being illegal immigrants.
That is caused by reckless immigration policy stoking resentment
It’s no longer more than a million a year (thanks Kier Starmer, I guess) but there are now at least 15 million people in the UK who aren’t White British, it is absurd to pretend that isn’t a big number.
A lot of people ideally want immigration to be negative, so +170k is still going in the wrong direction.
There are potential arguments in favour of that level of immigration but you need to convince people that it’s fine for their country to change beyond all recognition in the space of a single generation.
And unsurprisingly, the people trying to do that are finding it a tough sell
Reform take peoples concerns about the massive increase in population numbers and the knock on effects of housing and infrastructure, and cultural clash since we swapped eu immigration for sub Saharan and South Asian, and point it to the small boats.
I have no problem with refugees. I do have problem with a million people a year arriving through Heathrow. Farage and co want that though as it’s good for business (more workers, higher asset prices) so they blame it on the tiny numbers of small boats.
Brexit meant we swapped immigration of mostly like minded Europeans and temporary workers for permanent immigration from poorer countries with more medieval views.
I’m fairly left on the spectrum, I was door stepping for remain, and for Lib Dems in 2017,2019 and a by-election, I’m very much on the tax wealth not work side, and I have no problem in paying tax despite earning more than 90% of the people in the county, including the old folk who just keep taking and never paying.
But it seems I’m not left enough because I have a problem with the massive amounts of legal immigration under a populist right wing government and I should be sent to a gulag.
In fact I would argue: if someone is truly left wing in the traditional sense (i.e. they believe Marx when he says politics is all about class conflict, they want to live in a society where workers have more power, they support trade unions etc.) then logically... they should oppose mass immigration.
Mass immigration exists because it's in the interests of the capital owning classes. It massively reduces the bargaining power of labour, and the potential power of trade unions.
Yet most supposedly "left wing" people support mass immigration and for some reason they can't understand they're voluntarily acting as willing foot soldiers for people they profess to hate
"anyone claiming the UK’s legal migration is any sort of problem right now, or in the past, should probably not be listened to in any way shape or form for anything they have to say on the matter."
This sort of position is, IMO, disgusting and immoral. I cannot express how appalled I am that someone would state something like this. Anyone who's a citizen of the UK most certainly has a right to express their opinion on legal immigration, either in favor or against.
I find this recommendation funny, because I completely agree from a sysadmin perspective.
I used to use Gitlab at work, and small teams would run into so many footguns with CI that we had to throw up guardrails to prevent mistakes. Far too many links to gitlab issues that were not fixed even after >8+ years of being open ended up biting us. With GHA I haven't had that experience, and same for all of my self hosted Forgejo instances. I used to hate using GHA from about ~2018 to 2021, but they've fixed a lot of things I disliked since then.
What do you need real anchors for? Sharing of pipelines?
Maybe I’ve been burned too much by pipeline maintenance (because we didn’t have yaml anchors?) but I rather have builds defined in make /bazel/etc than in yaml. So the only thing the pipeline does is optionally restoring caches, kicking off the build system, uploading PR validation results, and saving cache. Pushing artifact etc is all done from inside the build system.
There is no “setup” like installing packages because we make the build image seperately.
i feel the same way. there's nothing interesting or curious about someone writing prompts to outsource a bunch of effort/thinking/work off to some most likely continuation sequence predictor / set of most likely continuation sequence predictors.
i miss seeing clever and thrifty solutions to weird problems. hell, i miss people just doing plain weird shit and posting that.
a first look through the comments on some first page threads and they seem far more thoughtful and considered. it generally seems like a nicer commenting environment that might actually do me some good moving forward.