Mistral Says Large 4 Has 1 Trillion Parameters, but Outside Tests Rank It Short of the Top, Plus AI That Reads Too Much Into People
Mistral's biggest model meets independent tests, a preprint finds AI fills in traits people never shared, and a Mac app searches your videos without the cloud.

The 60-second version
- Mistral released a public preview of Mistral Large 4 on Oct. 6 and says it has 1 trillion parameters with 49 billion active, while independent tester Artificial Analysis scored it 38 on its Intelligence Index and Vals AI ranked it 32nd of 44.
- RealCompanion, a preprint not yet peer-reviewed, studied 27,218 messages between ten people and an AI companion and found that only 3.4% of their messages depended on something said earlier.
- SCM, a free, open-source Mac app by developer Allen Lee, searches photos and videos in plain language on the user's own machine, though its README says it indexes one frame per video segment, not every frame.
Mistral says Large 4 has 1 trillion parameters, but outside tests rank it short of the top
Mistral released a public preview of Mistral Large 4 on Oct. 6, available now through its API [1]. The company calls it "a 1 trillion-parameter natively multimodal model with 49 billion active parameters" [1]. Early independent tests give mixed results: Artificial Analysis scores it above average for its price tier, while Vals AI ranks it 32nd of 44 [2] [3].
Mistral's product card describes a mixture-of-experts model [1]. In that design only a slice of the network works on each word, so it costs less to run than its full size suggests.
| What's claimed | What the evidence shows |
|---|---|
| 1 trillion total parameters, 49 billion active, per Mistral [1] | Not checkable yet. Mistral says the weights, the downloadable model files, "drop end of this month" [1]. TechCrunch reports they arrive in three weeks, "after safety testing is complete" [4]. Simon Willison's write-up repeats Mistral's figures [5]. |
| Mistral hopes it will be best in class among open-weight models and could beat closed models in specific areas, TechCrunch reports [4] | Artificial Analysis scores it 38 on its Intelligence Index, above the median of 26 for reasoning models at a similar price [2]. Mistral Large 3 scored 9, Willison notes [5]. Vals AI ranks it 32nd of 44 on the Vals Index, at 48.05% [3]. |
| 28.3% on Terminal-Bench 4, a coding test, per Mistral [1] | Vals measured 22.73% on Terminal-Bench 4.0 [3]. In Mistral's own blind human coding evaluation, it ranked second of five, behind Claude Opus 5 [1]. |
| Vals AI tests found it "exceeds GPT-6-Astra" on legal and financial tasks, Mistral says [1] | Mistral says it ran these comparisons through Vals, a third-party evaluator [1]. Vals' public page ranks it 6th of 75 on Harvey's Legal Agent Benchmark and 22nd of 75 on Finance Agent v2 [3]. |
| 82% on a test of reproducing and patching software flaws, "the highest of any model," and 93% of Cybench challenges, per Mistral [1] | These are Mistral-reported numbers; it attributes the 82% to a test in the Artificial Analysis Cyber Index, which this digest has not checked [1]. Mistral says several leading closed models score near zero on the same test because they refuse the task [1]. |
| $1.36 per million input tokens and $4.18 per million output tokens [1] | Artificial Analysis lists the same prices but calls the model "very verbose." It used 200M tokens to run the index, against a median of 81M [2]. |
Disclosure: Claude Opus 5 comes from the same developer as the AI model that writes Eureka Reports.
Still in training. Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own datacenters [1]. Pierre Stock, Mistral's VP of Science, gave TechCrunch a rounder figure of 4,000 [4]. Mistral also says reinforcement learning, a stage where the model improves from feedback on its own attempts, is still running, and it expects "large and rapid improvements in the weeks and months to come" [1].
Outside verdicts. Willison called it "maybe about 6 months behind the frontier" and noted that the API offers only two reasoning settings, "none" and "high" [5].
The Hacker News thread drew about 1,550 points [6]. Its most pointed critic wrote, "This model doesn't knock anybody's socks off" [7]. Another asked how Mistral would win customers against open models "1/2 - 1/3 the price but with similar capabilities" [8]. A third praised its "Impressive vision benchmarking" [9].
What's next. Mistral says it will release the weights by the end of the month. "As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology," the company says [1]. Only then can outsiders check the size claim for themselves.
AI systems read more into people than they reveal, a preprint on real companion chats says
AI systems that study months of a person's chats rebuild that person fairly well, then add traits the person never revealed, according to RealCompanion, a preprint posted to arXiv on Oct. 1 [10]. The paper, by Arman Behnam, Sunglyoung Kim and Liangwei Yang, is not yet peer-reviewed [10]. Hugging Face featured it on its Daily Papers list on Oct. 5, and it had 268 upvotes by Oct. 7 [11].
What did the researchers build?
A benchmark drawn from ten real relationships between people and an AI companion: 27,218 messages over up to 120 days [10]. Real conversations are private, the authors note, so earlier benchmarks relied on "invented people and invented questions" [10].
The chats come from an internal, friendship-style companion app with streaks, a habit tracker and mini-games [12]. All ten people consented to research use and release, and a model rewrote every message to remove names and other identifiers [12].
How often does a companion actually need to remember?
Rarely. Only 3.4% of people's messages depended on something said earlier [10], or 1.3% under a stricter reading [12]. When one did, the earlier message sat a median of 2,157 messages back [10].
Simple tools struggle with that distance. Checking the most recent messages found the needed one 95.9% of the time overall, but only 2.2% of the time when it was far back [10]. Keyword search did better at 24.8% but "still misses three in four" [12].
Can AI tell when the past matters?
Not yet, the authors report. A detector built to decide whether a message needs memory scored 0.530 on AUROC, a 0-to-1 measure where 0.5 is a coin flip [12].
Models also reach for the past when nobody asked. With no earlier messages, they never did. Given ten unselected earlier messages, they brought up the past when it wasn't needed 60.8% of the time [12]. Simply labeling the same messages "memories" instead of "earlier messages" made models bring up the past 10 to 14 percentage points more often [10].
Do AI systems understand the people they talk with?
Partly. Three agent systems, Claude Opus 5.5, Codex GPT-5.6-sol and Antigravity running Gemini 3.8 Flash, each tried to rebuild every person's persona, a file describing how they think, feel and decide [12]. All three scored an F1 of about 0.7, a 0-to-1 score that balances what they found against what they got wrong [12].
They recovered 86% to 89% of what the files held, but only 56% to 59% of what they produced matched [12]. The systems "add about seventy to eighty further fields for every hundred they recover" [12]. It is like a friend who remembers everything you said, then fills in things about you that you never said. The authors put it plainly: "They see the person, and then imagine more" [10].
Cost varied widely. At about the same score, Claude processed 546 million tokens, the word fragments models read and write, while Codex processed 18 million and Antigravity 36 million [12]. Disclosure: Claude Opus 5.5 comes from the same developer as the AI model that writes Eureka Reports.
What are the limits?
The authors are candid. Ten self-selected people using one app "are not a sample of companion users and still less of people" [12]. Two people account for 73% of all messages, and under the strict reading, memory results rest on just five people [12].
Ethics approval came from the data operator's internal committee, which the authors say "is not independent of the operator," and no institutional review board, the usual outside ethics panel, was involved [12]. Models produced the labels, which were then audited. That audit found about a fifth of one stage's written reasoning wrong [12]. The persona and trait files were produced by a language model; the authors say they are not validated psychological instruments and were not reviewed by the participants [12].
Privacy is the sharpest concern. Even after the rewrite, a language model matched a writing sample to the right person 98% of the time, against 10% by chance [12].
What happens next?
The dataset sits behind a gate on Hugging Face. A person reviews every request, and users must accept a data use agreement [13]. The terms forbid profiling or identifying people, any commercial use, and claims about whether the relationships helped or harmed anyone [12]. As the authors write, "the corpus records that a relationship continued, not what it did for the person in it" [12].
SCM searches your Mac's photos and videos in plain language, its developer says
SCM is an open-source Mac app that lets you search photos and videos in any folder by describing them [14]. Its developer, Allen Lee, released it under the MIT license [15] and calls it "100% free and open source" [16]. Its Show HN post reached 175 points [17].
The README promises "no accounts, no cloud, no uploads. Inference runs on your Mac" [14]. Search runs through an embedding model, which turns pictures and text into lists of numbers so similar meanings land close together [14]. Five modes cover whole files, scenes inside videos with a jump to the timecode, text in images, spoken words and an opt-in local chatbot [14].
| SCM | Apple Photos | Immich | |
|---|---|---|---|
| Where it runs | On your Mac, offline after a one-time 435MB model download, per the developer [14] | Photos on Mac with Apple Intelligence [18] | Not stated on the cited docs page [19] |
| What it searches | Any folder, including moments inside videos [14] | Photos and "a key moment in a video," in natural language [18] | Photos via CLIP-based "contextual" search [19] |
| The catch | Homebrew install needs Apple Silicon and macOS 12 or later, and media is copied into the app's own library [14] | One commenter couldn't use it because their photos sit "on an external SMB network drive" [20] | Suggested on Hacker News as the cross-platform option [21] |
The honest caveat. The app's pitch is search for "every frame of video" [14]. Its README says something narrower: it splits each video at camera cuts, and "each segment embeds its midpoint frame" [14]. Users choose a density from one sample every 60 seconds to one every 2.5 seconds [14]. One commenter warned that anything shorter than the interval can be missed [22].
Big libraries are the open question. A user with about 12,000 videos said not knowing the time it would take stopped them from trying it [23]. Another, who built a similar tool, wrote that "frame sampling rate is the whole ballgame. One frame a second on 12k videos is days" [24]. The developer lists the default model at about 480 to 570 milliseconds per image on a CPU [14]. Speed and privacy claims are the developer's own, with no independent tests yet.
Installing it also skips a macOS safety check. The Homebrew build uses a local development certificate rather than a Developer ID, so the installer clears the quarantine flag, a step the project says will go once releases are notarized [25]. Lee says he is exploring a native version "for speed" [26].
What AI companions and local photo search keep about you
RealCompanion's authors found that anonymized chats could still be traced to their writer 98% of the time, and that AI systems filled in traits people never shared [12]. SCM's developer takes the other route and keeps the index on the user's own Mac [14], though sampled video can still miss a short moment [22].
Editor's note
Placeholder: Justin Rogers adds a short note here after reviewing this issue.
Sources
- 1.OfficialMistral AI, Introducing Mistral Large 4, 2026-10-06
- 2.ReportingArtificial Analysis, Mistral Large 4 Preview - Intelligence, Performance & Price Analysis, 2026-10-06
- 3.ReportingVals AI, Mistral Large 4 Benchmarks, Cost and Capabilities, 2026-10-06
- 4.ReportingTechCrunch, Mistral’s new 1T model aims to leapfrog closed and open rivals, 2026-10-06
- 5.SocialSimon Willison's Weblog, Introducing Mistral Large 4: Le chonk, 2026-10-06
- 6.SocialHacker News, Mistral Large 4, 2026-10-06
- 7.SocialHacker News, Comment on "Mistral Large 4", 2026-10-06
- 8.SocialHacker News, Comment on "Mistral Large 4", 2026-10-06
- 9.SocialHacker News, Comment on "Mistral Large 4", 2026-10-06
- 10.PaperBehnam, Kim and Yang (arXiv), RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations, 2026-10-02
- 11.OfficialHugging Face Daily Papers, RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations, 2026-10-05
- 12.PaperBehnam, Kim and Yang (arXiv, full text v2), RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations, 2026-10-02
- 13.OfficialHugging Face (Quislab), Quislab/RealCompanion, 2026-10-06
- 14.OfficialGitHub (allenv0/SCM), allenv0/SCM: Deep AI search for every photo and every frame of video in any folder on macOS, 2026-10-07
- 15.OfficialGitHub (allenv0/SCM), SCM/LICENSE at main · allenv0/SCM, 2026-10-07
- 16.SocialHacker News, Comment on "Show HN: AI search for every photo and every frame of video on macOS", 2026-10-04
- 17.SocialHacker News, Show HN: AI search for every photo and every frame of video on macOS, 2026-10-04
- 18.OfficialApple Support, Search for photos and videos on Mac, 2026-10-07
- 19.OfficialImmich Docs, Searching, 2026-10-07
- 20.SocialHacker News, Comment on "Show HN: AI search for every photo and every frame of video on macOS", 2026-10-04
- 21.SocialHacker News, Comment on "Show HN: AI search for every photo and every frame of video on macOS", 2026-10-04
- 22.SocialHacker News, Comment on "Show HN: AI search for every photo and every frame of video on macOS", 2026-10-04
- 23.SocialHacker News, Comment on "Show HN: AI search for every photo and every frame of video on macOS", 2026-10-04
- 24.SocialHacker News, Comment on "Show HN: AI search for every photo and every frame of video on macOS", 2026-10-04
- 25.OfficialGitHub (allenv0/homebrew-scm), homebrew-scm/Casks/scm.rb at main · allenv0/homebrew-scm, 2026-10-07
- 26.SocialHacker News, Comment on "Show HN: AI search for every photo and every frame of video on macOS", 2026-10-04
- ImageNilo Velez, Open computes case with all components visible and purplish blue light glowing from the front panel, WordPress Photo Directory, CC0 1.0