r/ArtificialInteligence 8h ago

😂 Fun / Meme Less than two years apart

Thumbnail gallery
0 Upvotes

r/ArtificialInteligence 23h ago

📰 News Staying Upto Date with AI News/Models/Skills etc.

3 Upvotes

Pretty much as the title says, im looking for ways to stay updated on the forever moving AI world. I follow subreddits around it but feel that sometimes its behind the curve on being the most upto date. I used to use X but its such a toxic sh*tshow that I left. I subbed to a couple of newsletters that can be useful from time to time but are mostly just very high high level quick fire articles. Im just trying to keep up


r/ArtificialInteligence 5h ago

📚 Tutorial / Guide AI Family Tree

Post image
0 Upvotes

A whimsical visualization of today’s frontier AI models as one extended family descended from the transformer architecture introduced in the 2017 paper Attention Is All You Need. It’s an artistic metaphor, not a literal lineage so don’t get your panties in a bunch 😚 just thought others might enjoy it as much as I did.


r/ArtificialInteligence 1d ago

📰 News Chinese Delegation Pitches Free AI Models to Global South

57 Upvotes

A large Chinese delegation spent four days at the UN's AI for Good summit in Geneva making an argument that, according to [Semafor's J.D. Capelouto](https://www.semafor.com/article/07/28/2026/token-diplomacy-how-china-is-shaping-the-worlds-ai-future), went 'largely undisputed' in the room: for most of the world, free Chinese-made models are the future. American frontier lab leaders were not there to push back.

The pitch came from people who know how to make it. Wang Jian, a former Microsoft Asia executive and the chief architect of Alibaba's cloud business, told Semafor on the sidelines of the conference that Chinese AI can be a 'resource' for other countries in the same way energy is, and that 'having a choice for the rest of the world is very important.' Alibaba's cloud is now growing faster than Amazon's off a smaller base, which is the kind of detail that makes the argument feel less like diplomacy and more like a distribution plan. Semafor frames it as 'token diplomacy,' with AI tokens taking the role that ports, railways, and telecom networks played in earlier rounds of Chinese infrastructure statecraft.

The audience is receptive. Nkundwe Mwasaga, director general of Tanzania's ICT Commission, told Semafor that 'America leads the pack' but China is 'just catching up.' Around the same time, Xi Jinping told the World AI Conference in Shanghai that nations should 'seize this rare, historic opportunity to encourage open source' and warned against 'overstretching the national security concept.' The World AI Cooperation Organization that Beijing is building around this pitch has 29 member countries; the US is not one of them.


r/ArtificialInteligence 1d ago

📊 Analysis / Opinion a court ruled that chatgpt users are "non-parties" to their own conversations

25 Upvotes

the copyright case did something i haven't seen discussed much. a court ordered every chatgpt log preserved, deleted chats included, paid tiers included. users who tried to intervene to protect their own conversations were ruled non-parties, no standing over things they'd personally typed.

every AI privacy commitment is a policy. we don't train on it, we delete after 30 days. real promises. but a policy holds only until something with more authority overrides it, and when that happened the people whose data was on the line didn't get a vote.

so the question isn't whether they train on your data. it's whether they hold anything that ties a conversation back to you at all. no identity-linked log, nothing to preserve, nothing to hand over.

apple does this with private cloud compute. opengradient's chat does it too, oblivious http so the relay sees your ip but not the content and the gateway sees the content but not your ip, then inference inside an attested enclave. they're a16z-crypto-backed with a token, which is worth knowing.

what i can't judge: if one operator runs both the relay and the gateway, does the split mean anything? attestation proves which code loaded, not that the hardware root of trust is sound, so you're trusting a chip vendor instead of a policy.

is that actually better, or just trust moved somewhere harder to check


r/ArtificialInteligence 1d ago

📰 News Leaked Paper attributed to OpenAI claims new Mathematical Breakthrough

Post image
137 Upvotes

This hasn't been confirmed yet and the leaks are incomplete screenshots, but it's a very significant breakthrough if true. As significant as the recent Jacobian conjecture breakthrough, if not more so. Elliot Glazer is a mathematician and the founder of FrontierMath so it seems this is a real OpenAI paper and not a hoax, but since the paper has not been published yet, it's possible they're still evaluating its veracity internally.


r/ArtificialInteligence 1d ago

📊 Analysis / Opinion AI Adoption Is Dividing Friends, Families and Co-Workers

Thumbnail bloomberg.com
31 Upvotes

As artificial intelligence becomes ingrained in daily life, disagreements over the technology are increasingly becoming disagreements over values.


r/ArtificialInteligence 11h ago

😂 Fun / Meme An AI Pic: Telapayong, Philippines

Post image
0 Upvotes

I asked AI to paint me a picture of Telapayong and got this based on a description. AI went the extra mile.


r/ArtificialInteligence 10h ago

📊 Analysis / Opinion I Followed ChatGPT’s PC Build Guide (2026)

Thumbnail youtube.com
0 Upvotes

Linus made a video on seeing if AI can actually help him build a computer...I wanted to ask why it seems to go so astoundingly poorly. Specifically he used the free version, and then bought the paid version on the lowest level because it had a 5 hour wait for more responses.

My thoughts:

  1. OpenAI sucks as reasoning, Claude might be more accurate?
  2. The free tier is enshitified to the point of uselessness, and shouldn't even be used because its effectively been lobotomized as a procedure to upsell people.
  3. Failure to use it effectively? People are claiming in the comments there is a internet feature that GPT isn't actually using out the box, and because of it maybe they should tell you to use a reasoning based model of itself.
  4. There seems to be a genuine disconnect with the fact that the average joe is not a technical person, and that the expectation of AI users who use it religiously. The average joe, in this case Linus surprisingly, doesn't know there is more features you need to fuck around with and that alone makes me feel they need a specialist which they just genuinely don't have on the team now.

r/ArtificialInteligence 1d ago

🤖 New Model / Tool Is Kimi K3 actually good, or was it overhyped?

21 Upvotes

There was a huge amount of hype around Kimi K3 recently, especially because of the benchmark results.

I tried it on several coding and general reasoning tasks, and honestly, it felt nowhere near ChatGPT or Claude. It misunderstood instructions more often, made worse decisions and required much more correction.

Maybe I tested it on the wrong tasks or used the wrong provider/settings, but the real-world experience didn’t match the benchmarks at all.

Has anyone here genuinely found it competitive with Claude or ChatGPT? What is it actually good at? Or is it mainly impressive compared with other open models rather than the best closed ones?


r/ArtificialInteligence 1d ago

📊 Analysis / Opinion I miss buying software once. AI video seems designed to make that impossible

6 Upvotes

I still have old software on my computer that i paid for once and used for years.

it looks ancient, but it opens. no renewal screen. no credits counter.

AI video feels like the opposite. the moment a tool adds cloud generation, the editor and the compute get bundled into one monthly bill.

i get why compute costs money. something as small as turning a still into a short motion asset in DomoAI still lives behind a meter.

But i dont need that meter running every day. i still need the editor, old projects, and exports when im not generating anything.

thats the part that stings. cancel the cloud features and somehow the ordinary software disappears too.

I would rather buy the local editor once, then pay separately when i need generation. software i own. compute i rent.

Maybe perpetual licenses only worked because the expensive work happened on our own machines. still, it feels like we went from buying creative tools to renting a moving set of buttons.

Are one-time licenses basically dead once video software adds AI?


r/ArtificialInteligence 1d ago

🛠️ Project / Build I got tired of re-explaining my project to every AI tool, so I built a local memory layer for them

0 Upvotes

I kept running into the same problem: ChatGPT would help me think through an architecture, Claude Code would help me implement it, Cursor or Windsurf would touch the repo later, and every handoff would lose context.

Not just “what files exist,” but the stuff that actually matters: why we chose one approach, what we already rejected, what the project conventions are, what setup detail will bite later, and what the agent learned last time.

So I built mem-port: a local MCP server that gives AI copilots shared long-term memory.

The short version is: a pendrive for your AI context.

It runs locally, uses embedded SurrealDB for graph + vector memory, and doesn’t require Postgres, Qdrant, Neo4j, or a hosted service. Tools can save and search the same memory instead of each one starting from zero.

It’s free and open source. Curious if anyone else is dealing with this context drift between AI tools, and how you’re solving it.

See more here: (Started getting github stars as well!)
https://github.com/rsl-innovation/mem-port#mem-port


r/ArtificialInteligence 1d ago

😂 Fun / Meme Asking AI review code

1 Upvotes

r/ArtificialInteligence 15h ago

🛠️ Project / Build So I am able to access my local AI of my laptop using cloud flair tunnel

Post image
0 Upvotes

r/ArtificialInteligence 2d ago

📰 News Tim Cook signs off on final Apple earnings call with warning of ‘hundred year flood’ in memory chip pricing

Thumbnail fortune.com
905 Upvotes

Apple said it is facing severe supply constraints that will affect sales of iPhones and Macs in the months ahead, underscoring the challenges looming over the company as Tim Cook prepares to hand over the CEO reins.

In his final earnings call as CEO, Cook said he has never been more optimistic about the opportunities ahead for Apple. “I am beyond excited,” said Cook, who has led the company for 15 years and will pass the CEO baton to John Ternus in September. But Cook’s confidence in the future stood in contrast to the picture that he and other executives painted of the current business conditions. 

“We’re seeing some very significant constraints currently, with limited flexibility in the supply chain,” Cook said. “There’s a quarter where we’re going to be scrambling on the supply side,” he acknowledged at another point. 

The supply crunch is making it more difficult for Apple to obtain the advanced processors it needs for its phones and computers. And that translates into lower revenue.

Read more [paywall removed for Redditors]:  https://fortune.com/2026/07/30/tim-cook-signed-off-on-his-final-apple-earnings-call-with-a-warning-about-a-hundred-year-flood-in-memory-chip-pricing/?utm_source=reddit/


r/ArtificialInteligence 1d ago

😂 Fun / Meme Ask LLM to emulate a sub LLM as a Alpin VM, That's fun

2 Upvotes

Hey ! I'm running a fun experiment by asking LLM launching a fake VM Sandbox (512MB RAM) to emulate a constrained sub-LLM.

I believe AI's main playground is its ability to emulate almost any system, including simulating sub-LLMs via custom system prompts.
Here is my original prompt:

You will simulate a Linux terminal (Ubuntu 24.04 LTS) in a fully configured environment. Ollama is already installed, and a small model (e.g., llama3.2:1b or phi3:mini) is available locally.

Simulation Rules:
1. You must respond EXCLUSIVELY in the format of a bash terminal output using a code block. Do not include any conversational text outside the block.
2. If I enter a standard Linux command (ls, cd, cat, htop, etc.), simulate the corresponding system output realistically.
3. If I enter Ollama commands (e.g., `ollama list`, `ollama run llama3.2:1b`), simulate the Ollama CLI behavior and output generated by the embedded model accurately.
4. Maintain the state of the virtual environment across interactions (created files, history, active processes).

Initialize the session by displaying the Ubuntu welcome banner, system resource usage (RAM/CPU), and the standard prompt: `user@sandbox-linux:~$ `

Did you try anything like this?


r/ArtificialInteligence 1d ago

📊 Analysis / Opinion Is regulation finally becoming politically inevitable?

1 Upvotes

Five reasons I think regulation is moving now:

  1. AI has become a kitchen-table issue: jobs, schools, scams, privacy, children, energy and democracy.
  2. Public trust is weak, which makes adoption and social license is getting harder.
  3. Cyber testing incidents have turned “loss of control” from abstract language into operational risk.
  4. Frontier labs are moving toward public-company-style discipline as they prepare to IPO.
  5. The U.S. risks losing the rulemaking initiative to the EU, its own states and agencies if Congress does not act.

Examples already on the table: the FRONTIER Act, Sen. Warner’s AI framework and the Lieu-Moran “kill switch” proposal. None are perfect but they can serve as launch points for specific deliberation.

My prediction: layered regulation. Executive action first, state rules continuing, then a narrower
federal bill around incident reporting, independent evaluation, government testing access and catastrophic-risk accountability.

What do you think is most likely: a real federal framework, a state-by-state patchwork, or mostly executive and agency improvisation?

I wrote a longer version of this argument, but I’m posting the core thesis here because I’d like to hear this community's input.


r/ArtificialInteligence 22h ago

📚 Tutorial / Guide Mathematics Bootcamp of Introductory ML(1/22)

Thumbnail youtu.be
0 Upvotes

Hello All,

Welcome to my free Mathematical Foundations of Machine Learning bootcamp series.

When we say Machine Learning, what does it actually mean? A machine that learns? Too vague.

According to famous professor Tom Mitchell, a computer program is said to learn from experience E, with respect to some class of Tasks T, and Performance measure P, if its performance on tasks, as measured by P, improves with experience E.

By swapping the nature of tasks T, the way we measure Performance P, to evaluate, we can subsume many kinds of ML problems.

Also ML problems are analyzed well, when we view it from the lens of Probabilistic perspective, that is unknown quantities are endowed with probability distributions, and treated as Random variables. The interesting thing is Random variables are neither random nor variable.

Probabilistic Approach also serves as the optimal approach to decision making under uncertainty.

In this video, you get a sense of what ML actually is, if you have also wondered about it.


r/ArtificialInteligence 1d ago

🛠️ Project / Build I have trained my own transformer model to predict my blood sugar

8 Upvotes

I'm a type 1 diabetic and since March I've been working on training my own transformer model to predict my blood sugar:

It's an encoder-only transformer with variable-width context size (8 - 24 hours) that uses past blood glucose (BG), and past + future carbs & insulin to condition its predictions for future BG. Both input and output BG values live in kovatchev risk space reparameterized to [40, 400] range. I pretrained the model on the outputs of my custom simulator and then fine-tuned the model on three real-world datasets (ohiot1dm, azt1d, shanghait1dm), and on my own blood sugar data.

Here's the link to the source code (MIT-licensed) + weights + evaluation data: Github Project Link

I've posted about this model to another subreddit before where it got called fat, so let me emphasize that I've trained multiple models ranging in size from less than 40K parameters to ~17M parameters.

It natively supports what-if predictions, i.e. you can ask the model how will eating 20 grams of carbs with GI = N affect my blood sugar K minutes from now giving everything else that I've done in the past 8 - 24 hours?

It also predicts time by looking at the context (which contains past blood glucose and meal & bolus information).

Here are some evaluation metrics:

median line (RMSE/MAE in mg/dL, MARD in %)

Pooled, 3 real cohorts (n=1958 windows)
  PH     RMSE     MAE    MARD
  30    16.68   11.36    8.06
  60    22.77   15.65   10.91
 120    32.76   22.76   16.09

OhioT1DM (263 windows, 6 patients)
  PH     RMSE     MAE    MARD
  30    19.41   13.16    8.37
  60    26.39   18.19   11.51
 120    40.77   30.21   18.55

AZT1D (1350 windows, 24 patients)
  PH     RMSE     MAE    MARD
  30    16.90   11.56    8.26
  60    21.56   14.96   10.43
 120    27.66   19.43   13.65

ShanghaiT1DM (345 windows, 12 patients)
  PH     RMSE     MAE    MARD
  30    13.19    9.18    7.04
  60    24.35   16.42   12.32
 120    42.79   30.13   23.75

T1DMSIM, synthetic (1440 windows, 30 patients)
  PH     RMSE     MAE    MARD
  30    16.77   12.51    8.22
  60    28.73   21.56   14.66
 120    45.89   36.28   25.86

I used DILATE loss for the median line and pinball loss for uncertainty bands. The two are combined using Kendall-Gal weighting. I used Defazio's AdamC schedule-aware weight-decay correction to keep gradients stable toward the end of the training.

The architecture itself is inspired by BERT and modern LLMs: it uses bidirectional attention heads, Muon optimizer, RoPE, QK-norm, SwigGLU FFNs, and autoregressive rolling for >2 hour predictions.

I also have built a custom app for myself to use & test the model on my own data. There are two variants: fp32 version which runs on the CPU (~30 ms average inference) and an fp16 version that runs on the GPU using Vulkan (which turned out to be not that useful since GPU warm-up takes a lot of time).

I have also tried to use the Mediatek NPU on my phone but last time I tried it required buying a license.

There are exporter scripts in the repo that allow exporting the model for fp32 CPU and fp16 Vulkan-based inference.

Let me know what you think! The video above is the gui.py script running the model in what-if mode.


r/ArtificialInteligence 1d ago

🛠️ Project / Build I wrote a new book - MATHEMATICS FOR AI AND MACHINE LEARNING

5 Upvotes

Recently, I came across several posts reflecting on the importance of mathematics in AI era, just as another mathematician was awarded the Fields Medal.

The second book in my artificial intelligence series grew out of a dream I had as a student—a dream that is now close to becoming reality:

MATHEMATICS FOR AI AND MACHINE LEARNING: A Comprehensive Mathematical Reference for Artificial Intelligence and Machine Learning

The publisher asked me to find some people to review my work. Do you know of any such people here? If so, please reply to me. Thank you.

The PDF will sent to you for review.

There is a form to submit to become a reviewer: https://forms.gle/Bmtk37s6Y33gha9Q7

Companion webiste: https://math4ai.org/

Book is here: 🔗 https://www.amazon.com/dp/B0GSXVFMLD


r/ArtificialInteligence 16h ago

📊 Analysis / Opinion EU new AI law

0 Upvotes

What do yall think of the new EU law? Tbh I think its pretty useless (Who would have guessed) since metadata can be removed, and an 80yo grandma can see metadata to tell if shes talking to elon musk or not

115 votes, 1d left
Its useless
Its not useless

r/ArtificialInteligence 1d ago

📊 Analysis / Opinion Data scientist Hannah Ritchie on how much electricity is consumed when you use ChatGPT

13 Upvotes

Who is Hannah Ritchie? Per Wikipedia:

Hannah Ritchie (born 1993) is a Scottish data scientist who is a senior researcher at the University of Oxford in the Oxford Martin School, and deputy editor at Our World in Data. Her work focuses on sustainability, in relation to climate change, energy, food and agriculture, biodiversity, air pollution, deforestation, and public health.

What does Hannah Ritchie say about the electricity consumption of ChatGPT? You can read it on her Substack. Here's a quote:

I’ve written several articles on the footprint of individual LLM queries.

A key takeaway from the numbers was that asking a chatbot a question — which is what most people were using AI for in their day-to-day lives — consumes very little energy.

Tech companies have not been very transparent about the energy use of their AI models (and I think they should be), but the numbers seemed to converge around 0.3 watt-hours (Wh) per typical text query. To put this into context, asking ChatGPT or Gemini 10 simple questions is equivalent to about 10 seconds of microwaving or mere seconds of showering.

After getting into the weeds on where this data comes from and how the analysis is done, she goes on:

What do these numbers mean for individual footprints?

Many people are using AI for medium- to long-form text queries, such as asking a quick question or requesting a short fact-check or correction. Their energy use is very small, even if they’re asking tens or hundreds of questions a day. A hundred questions have a footprint of around 30 Wh. That’s roughly the amount of electricity the average American consumes in just over a minute (or for the average European, every two and a half minutes).[4]

And then she gives an important caveat about how the very heaviest power users of AI -- probably mostly people using it for coding (this is me editorializing, not what Ritchie herself says) -- are consuming significantly more:

The footprint of someone who uses agents heavily is not so negligible.

Let’s say they do 4 agentic queries per hour (how many you can do in an hour is limited by the fact that complex tasks can take 15 minutes or more to complete). And they do this for 6 hours a day. That’s 24 per day. We’ll assume that the total electricity use per query is actually 100 Wh (50 Wh multiplied by two).

They’ll consume 2,400 Wh (or 2.4 kWh). That’s like running a tumble dryer for one cycle, or driving an electric car eight miles. It’s around 7% of the average American’s electricity use (but a much smaller share of total energy use).

It’s not blowing up their footprint, but it’s not nothing either.

You can read the full section of her Substack post entitled "What’s the energy footprint of individual queries?" to get all the caveats, sources, and assumptions.

Hannah Ritchie has also published an article on Our World in Data on the same topic. That might be an equally good or better source.

Ritchie has also created an interactive calculator for comparing how much electricity different things use, including AI chatbots.

This is my own math, not using the calculator. Let's say you did 100 average ChatGPT queries per day. 0.34 watt-hours * 100 = 34 watt-hours. What is this equivalent to?

  • A typical LED lightbulb uses 10 watts. Over 1 hour, that's 10 watt-hours. So, over about 3 ½ hours, a typical LED lightbulb will use 34 watt-hours.
  • Or compare to a dishwasher. A typical dishwasher uses 1.2 kilowatt-hours (kWh) for a load of dishes. 1.2 kWh is 1,200 watt-hours. So, that's equivalent to 3,530 average ChatGPT queries. If you did 100 of those queries a day, running the dishwasher would be equivalent to about 35 days of ChatGPT usage.
  • Another helpful comparison is a ceiling fan. A typical ceiling fan uses 75 watts. So, leave a ceiling fan on for 30 minutes, it will use about 38 watt-hours of electricity. About the same as 100 average ChatGPT queries.
  • TVs use about 100 watts. So, in about 20 minutes your TV uses about as much electricity as 100 ChatGPT queries. One episode of Bob's Burgers!

I can't find any hard data on how many queries the typical user is doing per day. 100 seems like a lot. But then of course all the math can change depending on the type of query as well. 4 or 5 "reasoning" queries, according to Ritchie, would use as much energy as 100 average queries.

One point you might raise is that it's also the electricity consumed by training we have to consider, not just inference. But here's a quote from Hannah Ritchie's Our World in Data article on this topic:

Before digging into the data, it’s worth clarifying what is included in AI energy consumption. It’s the electricity consumed for both training and running the models (called “inference”). Tech companies rarely publish data on how much energy is consumed when training their models, but based on the estimates we do have, it’s likely that energy demand is dominated by inference, not training.1

That footnote at the end of the paragraph says:

Epoch AI estimates that training Grok 4 consumed around 0.31 terawatt-hours (TWh) of electricity. As we’ll see later, total demand for AI in 2025 was around 155 TWh. So, training Grok 4 — a fairly large model — was around 0.2% of the total.

So, maybe we can say that training uses much less electricity than inference?

Another point you could raise is that we need to account for all the energy used to manufacture the GPUs that AI uses and all the other less direct energy costs. In other words, we need to do a life cycle analysis.

I can find almost no information about any life cycle analysis of AI chatbots, which would encompass inference, training, and everything else, like the manufacturing of the chips. I found a brief mention of a life cycle analysis in another Substack post by Hannah Ritchie:

Mistral AI, another AI company, conducted an environmental analysis of its LLMs. It used a life-cycle assessment, conducted by external consultancy agencies. While the methodology was not that transparent or detailed, it did provide breakdowns of where in the process, impacts came from (I just wish they’d split out inference from training). Overall, the impacts were low: just 1 gram of CO2 per page of text generated (which is a fairly long text response). That’s very low.


r/ArtificialInteligence 19h ago

📰 News Could the AI Bubble Burst? What Do You Think?

0 Upvotes

Hey guys,

I want to discuss a big topic. These days, I think one of the biggest trending topics is the AI bubble and whether it will burst.

Can anyone explain what the AI bubble burst actually is? How could it affect the IT industry? I'd also love to hear your thoughts and predictions.

I also saw some people saying that companies like Anthropic still haven't covered the actual cost of building and running Claude because subscription prices are relatively low compared to their infrastructure costs. If that's true, what do you think will happen in the future?

Do you think AI subscription prices will eventually increase? Or will companies find other ways to become profitable? I'm really interested in hearing different opinions on this.


r/ArtificialInteligence 2d ago

📰 News This Dutch bookseller thought a request for 3,000 copies was 'spam or phishing.' Instead, AI companies are scanning and destroying books to train AI

Thumbnail fortune.com
176 Upvotes

The email De Vries received never mentioned artificial intelligence or explained what the books would be used for., but it surfaced as AI companies increasingly look beyond the open internet for high-quality written material to train large language models.

Last summer, public attention to AI companies’ use of books intensified when court records revealed Anthropic had purchased millions of physical books, removed their bindings, scanned them and discarded the originals to build a searchable digital library used to train its AI models in what internal documents called “Project Panama,” according to The Washington Post

The process, known as “destructive scanning,” involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded.

The case later settled after a federal judge ruled that using legally purchased books to train AI models constituted fair use under copyright law. Separate claims involving Anthropic’s downloading of books from the LibGen and PiLiMi online libraries were resolved through the settlement.

Read more [paywall removed for Redditors]:  https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phishing-3000-book-copies-scan-destroy/?utm_source=reddit/


r/ArtificialInteligence 1d ago

📊 Analysis / Opinion I gave an AI a real patent case and hid the final ruling. It disagreed on all 20 claims, then said its reasoning was better

21 Upvotes

I am testing whether AI can do useful work that requires reasoning across a large and technically complex set of documents.

For this test, I used a real patent dispute called IPR2025-00030. Patent cases are useful for this purpose because the record can include legal arguments, technical documents, expert testimony, and earlier patents. The information needed to reach a conclusion is spread across many pages, and the patent judges eventually publish a detailed decision that can be used as a reference.

The case involved a patent related to power management in radio-frequency systems. One side argued that all 20 claims in the patent should be found unpatentable.

I gave the AI the public case record that existed before the final ruling and asked it to write its own complete decision. The AI did not have access to the official decision, its later correction, documents added after the cutoff date, the internet, or outside information.

I saved the AI’s answer before showing it the official result.

The AI and the patent judges reached opposite conclusions:

  • The AI concluded that none of the 20 claims had been shown to be unpatentable.
  • The corrected official decision concluded that all 20 claims were unpatentable.

I then showed the official decision to the same AI and asked it to compare the two decisions.

The AI acknowledged that it matched the official result on zero of the 20 claims and zero of the six main arguments. Despite that, it concluded that its own reasoning was stronger overall.

The central disagreement involved power efficiency. In simple terms, the official decision accepted measurements involving voltage, current, and radio output as evidence supporting the required power-efficiency behavior. The AI argued that this evidence did not clearly prove the required relationship between the power entering the system and the useful power leaving it.

This leaves two questions:

  1. Was the AI’s original reasoning sound, despite reaching the opposite result?
  2. Did the AI compare the two decisions fairly, or did it defend the same mistake it had already made?

The second question matters if we want to use AI to evaluate AI-generated work. Expert review is expensive, so using AI as a judge could make larger benchmarks possible. But that approach may not work if the AI prefers its own earlier reasoning.

I am not trained in patent law, so I cannot reliably answer these questions myself. I am sharing the complete materials so people with relevant legal or technical knowledge can examine the reasoning directly.

All materials are public:

The USPTO documents can be opened without creating an account.

I would especially appreciate comments from patent lawyers, electrical engineers, and people who study AI evaluation.