this post was submitted on 27 Mar 2026

607 points (97.2% liked)

Technology

83220 readers

3079 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

607

Announcing ARC-AGI-3 - A benchmark that tests if AI can explore, learn, and adapt in unfamiliar situations. Humans score 100%. Frontier AI scores 0.26%. (arcprize.org)

submitted 4 days ago* (last edited 4 days ago) by brianpeiris@lemmy.ca to c/technology@lemmy.world

192 comments fedilink hide all child comments

The ARC Prize organization designs benchmarks which are specifically crafted to demonstrate tasks that humans complete easily, but are difficult for AIs like LLMs, "Reasoning" models, and Agentic frameworks.

ARC-AGI-3 is the first fully interactive benchmark in the ARC-AGI series. ARC-AGI-3 represents hundreds of original turn-based environments, each handcrafted by a team of human game designers. There are no instructions, no rules, and no stated goals. To succeed, an AI agent must explore each environment on its own, figure out how it works, discover what winning looks like, and carry what it learns forward across increasingly difficult levels.

Previous ARC-AGI benchmarks predicted and tracked major AI breakthroughs, from reasoning models to coding agents. ARC-AGI-3 points to what's next: the gap between AI that can follow instructions and AI that can genuinely explore, learn, and adapt in unfamiliar situations.

You can try the tasks yourself here: https://arcprize.org/arc-agi/3

Here is the current leaderboard for ARC-AGI 3, using state of the art models

OpenAI GPT-5.4 High - 0.3% success rate at $5.2K
Google Gemini 3.1 Pro - 0.2% success rate at $2.2K
Anthropic Opus 4.6 Max - 0.2% success rate at $8.9K
xAI Grok 4.20 Reasoning - 0.0% success rate $3.8K.

(Logarithmic cost on the horizontal axis. Note that the vertical scale goes from 0% to 3% in this graph. If human scores were included, they would be at 100%, at the cost of approximately $250.)

https://arcprize.org/leaderboard

Technical report: https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf

In order for an environment to be included in ARC-AGI-3, it needs to pass the minimum “easy for humans” threshold. Each environment was attempted by 10 people. Only environments that could be fully solved by at least two human participants (independently) were considered for inclusion in the public, semi-private and fully-private sets. Many environments were solved by six or more people. As a reminder, an environment is considered solved only if the test taker was able to complete all levels, upon seeing the environment for the very first time. As such, all ARC-AGI-3 environments are verified to be 100% solvable by humans with no prior task-specific training

you are viewing a single comment's thread
view the rest of the comments

[–] PhoenixDog@lemmy.world 6 points 2 days ago* (last edited 2 days ago) (1 children)

Someone else in the comments said it perfectly. AI is just data regurgitation. It's like calling me highly intelligent because I read you a paragraph from Wikipedia. I didn't know anything. I just read a thing and said it out loud.

[–] mechoman444@lemmy.world -2 points 2 days ago (3 children)

No. You're not just wrong, you're aggressively uninformed.

By you repeating the same tired “AI is just regurgitating data” line makes it clear you don’t understand what you’re criticizing. Calling large language models “AI” the way you are doing it just exposes that you do not know what you are talking about. It is like a creationist smugly saying “orangutang” instead of “orangutan” and thinking they sound informed. You are not demonstrating insight. You are advertising ignorance.

What you’re describing, reading a paragraph off Wikipedia, is literal retrieval. That is not how modern language models operate. They are not databases with a search bar attached. They are probabilistic systems trained to model patterns, structure, and relationships across massive datasets. When they generate a response, they are not pulling a stored paragraph. They are constructing output token by token based on learned representations.

If it were just regurgitation, you would constantly see verbatim copies of training data. You do not. What you see instead is synthesis. Concepts are recombined, abstracted, and adapted to context. The system can explain the same idea multiple ways, shift tone, handle novel prompts, and connect ideas that were never explicitly paired in the source material. That is fundamentally different from reading something out loud.

Your analogy fails because it assumes nothing is being transformed. In reality, transformation is the entire mechanism. Information is compressed into weights and then expanded into new outputs.

Is it human intelligence. No. Is it perfect. No. But reducing it to “just reading Wikipedia out loud” is not skepticism. It is a basic failure to understand how the technology works.

If you are going to criticize something, at least learn what it is first.

[–] lordbritishbusiness@lemmy.world 4 points 2 days ago (1 children)

Counterpoint: Why should they learn about it?

It is a good thing to reduce ignorance, but there is more to learn in the world than there is time to learn or space in the brain. People must specialise.

You must accept that not everyone will understand everything, and this is okay.

The nature of a Large Language Model is very specialist knowledge, data regurgitation is apt from a distance, especially when most publically available models are primarily used for search.

Criticism must be accepted, even from those who do not understand, so long as it's in good faith. It is after all an opportunity to reduce ignorance to someone with the time and interest to learn.

Don't rudely lord your intelligence over someone else, it might not end well, and invalidates the delivery of your entire argument.

[–] mechoman444@lemmy.world -1 points 2 days ago

The reason he should learn about it is because he's talking about it as though he's informed and he is not.

I don't have to be a LLM programmer working at openai to have a working knowledge of how these machines function. It's literally just a Google search.

He made an unreasonable ignorant comment and I called him out. He should feel ashamed and I have absolutely no reason to pad down what I'm saying under the guise of being nice.

[–] PhoenixDog@lemmy.world 3 points 2 days ago* (last edited 2 days ago) (1 children)

This might be the most comprehensive comment I've ever read about someone saying how utterly stupid they are to the world. It's incredibly impressive how articulate you described your absolute lack of critical thinking.

It's almost like intentionally shooting yourself in the nuts, and openly releasing the video of it saying you promote gun safety.

[–] mechoman444@lemmy.world -2 points 2 days ago

Calling an llm a Wikipedia regurgitator is factually and objectively incorrect.

Is there anything that you can say to refute the facts that I presented in my above comment?

(I rolled my eye so hard at your comment that I pulled my back out)

[–] hitmyspot@aussie.zone -1 points 2 days ago (1 children)

You're discounting the fact that a human reading Wikipedia will attribute intonation and tone to the text to give further context and meaning. I think the analogy is good. Its not precise but it is the same thing.

I do think AI has a useful purpose and is here to stay. I don't think it's groundbreaking like the AI companies want us to think. The bubble will burst and then we'll see where the cards lie.

OpenAI has lost their lead and I expect they will start to struggle with further funding. There are quite a few warning signs. The price of oil is likely to increase power prices generally and cause construction delays and cost rises. Both will hamper their plans. They still don't have a viable model for profit.

[–] mechoman444@lemmy.world -3 points 2 days ago (1 children)

The analogy is terrible and is not at all, once again, what llms do.

This is an objective fact I have provided evidence to support this.

How are you saying the analogy is good?

[–] hitmyspot@aussie.zone 3 points 2 days ago (1 children)

Ana analogy does not need to be precise. It expresses a comparison for easier understanding. It is not what LLMs do. However what you've expressed is simplified also. So by your standard, it is not useful for the discussion.

So maybe get your head out of your ass and try to understand what people are trying to express instead of correcting them when they are not incorrect.

If precision was of that much importance to you, you would have a different opinion of LLMs.

[–] mechoman444@lemmy.world -1 points 2 days ago (1 children)

I fully understand the analogy being presented. It is a poor analogy and fundamentally incorrect because that is not how LLMs function. They do not “read back Wikipedia pages,” which is a complete misunderstanding of the technology, not a minor lack of precision.

I am not disputing that it is an analogy, nor am I claiming that exact precision is necessary to analyze it. The point remains: the analogy fails.

What is curious is how people focus on my tone, saying I am aggressive or should be more precise, rather than engaging with the substance of my argument. So far, no one has directly refuted my points. This suggests that many responding are simply following the anti-AI bandwagon without understanding the technology, which is both reductive and disappointing.

[–] hitmyspot@aussie.zone 2 points 2 days ago (2 children)

No, the analogy is about not understanding but regurgitating data. It's more complex than that but the gist is that they don't understand or have knowledge of the data being presented.

They are statistical models for what is desirable output. They don't understand what they give as an answer. That is why they halluncinate information that sounds plausible and confident.

We're not refuting your point about how the technology works, but rather that the person you replied to provided a poor analogy. They didn't. It served the purpose it was designed to do. If you don't understand that, that's on you, not them. Maybe ask an ai to explain. ;)

[–] PhoenixDog@lemmy.world 1 points 1 day ago

Nailed it.

[–] mechoman444@lemmy.world 0 points 1 day ago* (last edited 1 day ago) (1 children)

Someone else in the comments said it perfectly. Al is just data regurgitation. It's like calling me highly intelligent because I read you a paragraph from Wikipedia. I didn't know anything. I just read a thing and said it out loud.

Christ on a stick.

The original analogy literally states "AI is just data regurgitation" now you're what? Saying it's more complex? Ever heard of a motte and Bailey. Cuz that's what you're doing now.

Once again, for the people in the back, the analogy is a failure. It does not work. Llms are not regurgitation machines.

Motte and bailey so it's faster for you to look up.

[–] hitmyspot@aussie.zone 1 points 1 day ago (1 children)

They simplified it, and also used hyperbole. That's not the same as motte and bailey.

You're being too literal, while still being imprecise. It's likely why you're struggling with what the analogy is for.

[–] mechoman444@lemmy.world 1 points 1 day ago (1 children)

He is claiming the analogy works, then retreating to a more defensible position by admitting the system is more complex.

I am not being overly simplistic or imprecise. I am stating plainly that the analogy fails. LLMs do not regurgitate stored information. They generate novel outputs by statistically modeling and interpreting patterns in their training data. I supported that position with objective facts, and no one has attempted to directly refute them. Instead, the responses rely on vague arguments about “precision” and “simplicity,” which do not address the core claim.

[–] hitmyspot@aussie.zone 1 points 1 day ago (1 children)

The analogy takes one sentence to explain. Your explanation took 5 paragraphs.

That is the point of an analogy. To make it "analagous" to something familiar, so as to avoid explanation.

You keep on asking someone to refute your facts. They are not in dispute. What's in dispute is whether the analogy works. It does as it represents the fact that the ai does not understand it's output. The analogy does not reference the search nor the human that is doing the reading.

The point is that the AI LLM does not understand what it is outputting, in the same way that a person does not need to understand a Wikipedia page to read it.

Rather than a Wikipedia page, perhaps you'd get the analogy better if they said reading a page from an advanced physics textbook. The point is that the information being presented accurately does not infer understanding in the case of AI. That was represented perfectly fine in the analogy, which was it's purpose.

[–] mechoman444@lemmy.world 0 points 22 hours ago

No. That is not what the analogy means. That is what you are choosing to extract from it because it supports the direction you want this exchange to go.

The use of the word “regurgitate” carries a very specific implication. It suggests that LLMs retrieve and repeat stored information verbatim. That is not how they function. We both appear to agree on that point.

LLMs do not rely on stored facts in the way the analogy implies. They generate outputs by modeling patterns in data, producing responses that are often novel rather than retrieved.

Whether or not the model understands or comprehends the content is irrelevant to this distinction. Comprehension is not a requirement for the system to function. So yes, the analogy is overly simplistic and ignores the actual mechanism at work.

To be precise: it does not matter that the model lacks awareness or understanding. It is still capable of analyzing patterns and generating new outputs from its training data. That is not regurgitation.

Concisely as I can: llms do not regurgitate data, the analogy fails.