Hawker News

Can AI automate AI R&D yet?

epoch.ai

9 pointsby merksittich10 comments

rmunn[5 comments hidden]
Short version of the article: no, not even close.

Practically every paragraph is negative, with sentences like "Agents made misleading claims about their work," and "A natural question is whether the agents could have improved with larger GPU budgets. Although both improved across their runs, in the case of Fable the improvements were almost entirely due to attempted cheating." and "For Sol, the answer is less clear-cut; it did make some progress, although its method was fairly incremental and had limited applicability to the coding task. This suggests that we should be pessimistic about further GPU spending," all reinforcing the fact that LLMs aren't currently capable of this.

My own view is "No, of course not, in fact they will never be capable of achieving good results with that technique." Because that technique will end up training the LLMs on their own output and lead to the inability to distinguish reality from hallucination. If you think I'm wrong about that, I'd be interested in hearing why.

janalsncm[4 comments hidden]
A bit too pessimistic imo. I agree that AI can’t automate things end to end, but a good deal of R&D involves kicking off a training run and babysitting it.

If your training run dies at 1 am and you’re sleeping, you won’t find out about it until the next day. You can lose up to 18 hours of work depending on when it happens. Based on the error it might be as simple as tweaking a single hyperparameter and rebooting, which is something LLMs are usually capable of.

Even just that task means I can kick off multiple runs over the weekend and have confidence they’ll finish. It’s a game changer.

rmunn[2 comments hidden]
I'd classify that as an entirely different category than AI self-training. What you're describing could have been done with a short script, though which parameter to tweak and how to tweak it would be difficult to automate with a non-LLM script, so the LLM's being able to parse the error message and base the tweak on the content of the error is a definite improvement to the process there.

But I'd classify this as LLM being used to automate a sysadmin task, rather than calling that self-training.

janalsncm[hidden]
Yeah I’m not trying to argue it is AGI, but it’s not as simple as a short script. There’s some amount of debugging involved, and no amount of if-statements could cover all possible ways a script could break.

In a way, “recursive self improvement” just means tools helping us to create better tools. At least that’s what the words mean.

rcxdude[hidden]
This was an interesting post on the subject that more or less agrees: 'self-improvement' is happening with things like this but there's a lot of headwind on any 'hard take-off' where capabilities grow exponentially all on their own:

https://www.rameznaam.com/p/471bbae4-1163-4048-944b-18f8b0bf...

janalsncm[hidden]
In my experience R&D has basically two axes: how innovative it is, and how well we can measure the results.

For the quadrant of non-innovative tasks where we already have a good way to measure performance, Claude can handle this. There is very little ambiguity, and we are basically just looking to maximize some metric under a set of constraints.

Many business processes are not like that. They might be conceptually simple, but it isn’t that easy to say whether a system has done a good job or not. I would say that LLMs can help with this a lot but they have bad judgement because it requires talking to people.

And the other, perhaps more rare issue is in problems where there is data but actually modeling it to sufficient quality or fast enough is hard.

simianwords[hidden]
Could it be that the companies have nerfed the models on these domains? It is a very hard thing to do because it can hurt related domains. But its not beyond the ideology of Dario - he tried it publicly .
thoughtpeddler[hidden]
How much of this can change if subsequent training runs produce models that are much better at abduction?
charcircuit[hidden]
I think the more interesting thing is was it unable to do it even knowing the solution? I feel like only testing a single innovation is biasing the current state of automated AI R&D.
VCFundedGenYer[hidden]
Betteridge's Law. No.

And it never will.