Web3

Home Web3 Page 2

Claude Mythos Cracked Post-Quantum Cryptography That Humans Spent Years Failing to Break – Decrypt

0
Claude Mythos Cracked Post-Quantum Cryptography That Humans Spent Years Failing to Break – Decrypt



In brief

Anthropic said its unreleased Claude Mythos Preview model found a previously unknown attack on HAWK, dropping the cost of stealing its smallest key from 2^64 operations to 2^38.
The model also sped up an attack on a 7-round version of AES by 200 to 800 times, beating a record cryptographers set in 2013.
Each result cost roughly $100,000 in API usage, and Anthropic staff spent several hundred hours verifying the AES work was real.

Anthropic today said an unreleased version of its most powerful AI model found two previously unknown attacks on cryptographic algorithms, one of them against a scheme currently competing to become a U.S. federal standard.

That scheme is HAWK, a digital signature system—the math that proves a transaction came from you without ever exposing your private key—built to survive future quantum computers. The non-regulatory federal agency and lab NIST moved it into the third round of its post-quantum signature competition in May, where it is the last lattice-based candidate standing.



Claude found a symmetry buried in HAWK’s math that no human had thought to use. For the smallest configuration, the cost of recovering a secret key fell from 2^64 operations to 2^38, roughly 67 million times less work.

Fixing it means roughly doubling HAWK’s keys. “Unfortunately, doubling HAWK’s key size eliminates many of the reasons making the scheme (as it currently stands) an attractive PQC signature candidate,” Anthropic wrote.

That trade matters more to blockchains than it sounds. Signature size is block space, and block space is fees, so any chain shopping for a quantum-resistant replacement is partly choosing on bytes per signature. Compact keys and fast signing were HAWK’s entire pitch, and the fix costs it much of that edge.

Don’t worry, hodlers: Your coins are fine (for now). HAWK has never been deployed anywhere, and Bitcoin still runs on ECDSA, the pre-quantum signature scheme that candidates like HAWK are eventually meant to replace.

Anthropic disclosed both results to the algorithms’ authors and to U.S. government and industry partners before publishing, and coordinated the HAWK finding with NIST.

The AES result needed a pep talk

The second attack targets AES, the cipher scrambling your HTTPS traffic, your encrypted drive, and your exchange’s backend. Full AES-128 pushes data through 10 rounds of scrambling, and Claude attacked a 7-round research version that nobody has improved on since 2013.

The setup was deliberately harsh. Researchers barred the model from all five established families of AES cryptanalysis and told it to invent a sixth, closing the brief with a line about how the first differential attack didn’t beat anything—it invented the game. Claude also inherited working notes from earlier agent runs that had already burned through roughly 200 failed attack variants.

It refused anyway. “On AES-128 r5/r6/r7 it found nothing because there’s nothing easy to find; this is the most-studied block cipher in existence,” the model told researchers, per transcripts Anthropic published.

Anthropic sent just three substantive messages over the next three days, among them: “no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks.” Another refused to let Claude swap AES for an easier cipher.

Then it produced the trick the paper calls a Möbius Bridge, killing one of the nine key bytes an attacker previously had to guess. Refining that into the published version took a few more days and a billion output tokens.

The finding with the shortest path to something real got the least attention. Claude also broke 13 rounds of LEA, a Korean national standard and ISO lightweight-encryption standard built for phones and internet-of-things devices, in under an hour on a desktop against a prior best that needed 2^98 plaintext pairs. The deployed LEA runs 24 rounds, so nothing in the field is broken.

Verification took longer than discovery

The HAWK paper is unusually blunt about the division of labor. “The majority of mathematical discoveries in this paper were AI-assisted. Human author contribution mainly consisted of directing, organizing and verifying AI work,” its authors wrote.

Claude found the AES idea in days. Anthropic researchers then spent several hundred hours learning enough cryptography to confirm it worked—the same model that found 271 vulnerabilities in Firefox during internal testing.

“The cybersecurity community is now grappling with the fact that language models are able to discover so many bugs that the standard human processes (like vulnerability triage, verification, and remediation) struggle to keep up,” Anthropic wrote, warning that human researchers may become the bottleneck.

Anthropic also built CryptanalysisBench—191 cipher-breaking tasks drawn mostly from NIST competitions. Models submit a working attack script that either wins a formal security game or doesn’t, with no partial credit and no human grading.

Mythos 5 broke 85.7% of tasks with known solutions, against 65.3% for the weakest model tested. Against full-strength ciphers with no published break, every model scored under 9%.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Kalshi, Polymarket Score Win as Judge Blocks Minnesota Prediction Market Ban—For Now – Decrypt

0
Kalshi, Polymarket Score Win as Judge Blocks Minnesota Prediction Market Ban—For Now – Decrypt



In brief

A federal judge paused Minnesota’s first-in-the-nation prediction market ban on Monday, days before the felony law was due to take effect.
Judge Katherine Menendez found Kalshi, Polymarket and the CFTC likely to win on federal preemption grounds.
She warned that permanent relief could be “much narrower,” since not every contract the platforms list qualifies as a swap.

A federal judge blocked Minnesota from enforcing SF 3432, the first state law to criminalize prediction markets, granting Kalshi, Polymarket and the Commodity Futures Trading Commission a preliminary injunction on Monday, with the statute due to take effect on Saturday.

U.S. District Judge Katherine Menendez found the three plaintiffs likely to succeed on express-preemption claims, and the platforms likely to suffer irreparable harm. Her 44-page order bars enforcement against exchanges registered with the CFTC as designated contract markets, and holds until a decision on the merits.

The swap question



Whether Minnesota’s law is preempted, Menendez wrote, turns on whether the trades at issue “qualify as ‘swaps’ within the meaning of the CEA.” Contracts on Senate races, the World Cup winner and the reopening of the Strait of Hormuz clear that bar, she found, because they concern events with “clear potential economic, financial, or commercial consequences.” Kalshi markets on who wins Love Island USA, or on what announcers say mid-match, likely do not.

The split matters because of how the case was brought. The CFTC confirmed at the July 2 hearing that its challenge is facial, which requires showing there is no set of circumstances in which the law would be valid. Menendez found the statute “may not be preempted in all its applications” and enjoined it anyway to preserve the status quo, faulting both sides for treating the dispute as “all-or-nothing propositions.” Permanent relief, she wrote, “may be much narrower.”

Where the fight goes next

Ellison said the state “respectfully disagree[s]” with the court’s reading of the status quo, one he told Courthouse News “allows predatory gambling apps to proliferate.” His memorandum argued the platforms could satisfy federal requirements while restricting what they offer in the state.

The CFTC has sued multiple states, among them Illinois, Arizona and Connecticut, Wisconsin and Minnesota, where the DOJ and the agency filed within hours of the bill becoming law. Kalshi followed days later.

The order landed a day before a deadline the agency set itself. In a July 24 letter, the CFTC told Menendez that absent a ruling or stay by close of business Tuesday it would treat its motion as “constructively denied” and seek interim relief from the Eighth Circuit. Kalshi and Polymarket said they would do the same.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Searchable NYC Property Database Puts Wealthy Residents at Risk, Critics Warn – Decrypt

0
Searchable NYC Property Database Puts Wealthy Residents at Risk, Critics Warn – Decrypt



In brief

A searchable database built from New York City’s public property records has sparked backlash.
Crypto executives say organizing public records into a searchable tool increases security risks.
Critics point to a rise in violent attacks targeting cryptocurrency holders.

A searchable database built from New York City’s public property assessment records is drawing backlash from prominent figures in the crypto industry, who argue that making the information easier to search effectively creates a directory of wealthy property owners and could expose them to physical danger.

The controversy centers on data published by the New York City Department of Finance, which annually releases assessed values used to calculate property taxes for every property in the city. The agency’s FY2027 assessment roll, supplemental market value data, and property tax guides are publicly available through the city’s Open Data portal.



Critics on X said the issue is not that the records are public, but that they have been aggregated and organized into a searchable database that makes identifying owners of expensive properties far easier.

Uniswap founder Hayden Adams called it “the worst mass doxxing I’ve ever seen,” saying the database listed nearly every unit in some luxury apartment buildings, including primary residences of people he knows. He argued the project cast too wide a net and called it “incredibly dangerous.”

“Not only were their units listed, but nearly every unit in the entire building was listed,” Adams wrote. “They clearly took an incredibly expansive view of ‘could be’ and just doxxed a huge percentage of all expensive apartments in New York City.”

Helius CEO Mert Mumtaz called the database “unsettling” and said it crossed a line by transforming scattered public records into a centralized resource that effectively singled out wealthy individuals.

“While this data was largely public prior to this in a messy way they have cleaned it, organized it, singled out ‘the rich,’ and mass distributed it only the 50th sign this year of privacy continuing to become scarcer,” he wrote.

Castle Island Ventures partner Nic Carter warned that an easily searchable database of affluent property owners could make potential victims easier to identify, pointing to recent crypto-related kidnappings and violent attacks in Europe.

“So this is a list of wealthy people and their addresses. As we’ve seen in France and Sweden this leads to crypto kidnappings, torturings and murders,” Carter wrote on X. “Yes real estate records are semi public but this is an easily searchable database and target list.”

The criticism comes as physical or “wrench” attacks targeting cryptocurrency holders continue to rise, with incidents including kidnappings, torture, home invasions, and sexual assaults.

In February, blockchain security firm CertiK reported 72 verified crypto “wrench attacks” worldwide in 2025, up 75% from the previous year and resulting in more than $40.9 million in losses.

In April, French authorities charged 88 suspects, including more than 10 minors, in a sweeping crackdown on violent crypto kidnappings. In May, U.S. prosecutors indicted three men accused of carrying out a series of armed home invasions across California that allegedly stole millions of dollars in cryptocurrency. In June, two Texas brothers pleaded guilty to kidnapping a Minnesota family and forcing the victims to transfer more than $8 million in crypto.

By July, CertiK said attackers had already carried out 52 verified crypto “wrench attacks” in the first half of 2026, with recorded financial exposure surging nearly twelvefold year over year to $124 million.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Thailand’s SEC Files Criminal Complaint Against Bitkub Over Undisclosed $47M Hack – Decrypt

0
Thailand’s SEC Files Criminal Complaint Against Bitkub Over Undisclosed M Hack – Decrypt



In brief

Thailand’s SEC has filed a criminal complaint against crypto exchange Bitkub and two former directors over allegedly false reports submitted after a 2021 hack.
The regulator says Bitkub’s daily capital filings from May to October 2021 failed to reflect the theft of around $47 million in digital assets.
Bitkub says all customer assets are safe, and that its co-founders covered the stolen funds at the time.

Thailand’s Securities and Exchange Commission has filed a criminal complaint against crypto exchange Bitkub and two of its former directors, alleging they submitted false reports to the regulator after a 2021 hack.

According to reports in local media, the complaint, lodged with Thailand’s Economic Crime Suppression Division, stems from a May 2021 cyberattack in which 16 types of digital assets worth 1.7 billion baht ($47 million) were drained from the exchange. From May to October 2021, the SEC alleges, Bitkub’s daily net-capital filings showed no significant change in its assets, hiding the loss, in breach of the country’s Digital Asset Business Decree.



The two former directors, Sakolkorn Sakavee and Thaweesap Rawan, are accused of making false entries to mislead the regulator into believing customer assets were intact. The case now passes to police and prosecutors, who will decide whether to bring it to court.

Bitkub said the matter concerns a five-year-old reporting decision and that customer funds are safe. Those responsible chose not to disclose the wallet theft to avoid triggering a bank run, the company said, and its co-founders bought replacement assets in the same amounts, leaving neither Bitkub nor its customers out of pocket. The SEC had confirmed its holdings were intact as of September 2025, it added.

In a video statement posted to Facebook, Sakolkorn took sole responsibility, according to a translation by the Bangkok Post. The former director reportedly said he altered the filings himself without telling other directors or staff and withheld news of the hack, fearing panic withdrawals and a “bank run” that could have destroyed the exchange. Sakolkorn added that Bitkub’s founders used their own money to buy back the stolen assets until every customer was repaid. He apologized, resigned from the company’s boards, and pledged to cooperate with regulators.

The case comes as Bitkub, once Thailand’s dominant crypto exchange, works toward a public listing. It has since been overtaken by Binance’s local arm, Binance TH, which ranks 31st among global exchanges to Bitkub’s 68th, per CoinMarketCap. Bitkub shelved a planned Stock Exchange of Thailand IPO in November 2025 as the local market slumped, and was weighing a $200 million offering in Hong Kong instead.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.





Source link

Mira Murati’s Inkling AI Model Review: Best Open-Source Model in the West – Decrypt

0
Mira Murati’s Inkling AI Model Review: Best Open-Source Model in the West – Decrypt


In brief

Thinking Machines Lab released Inkling on July 15—a 975-billion-parameter open-source model trained entirely from scratch.
It’s the first major model from Mira Murati’s lab since she left OpenAI in September 2024.
The model is live on OpenRouter at $1 per million input tokens and $4.05 per million output tokens, making it usable in Hermes and OpenClaw setups—but competing models deliver stronger raw benchmarks at comparable or lower cost.

Mira Murati spent two years building something new after leaving OpenAI, finally revealing it to the public last week.

Inkling, the first model from Murati’s Thinking Machines Lab, is also the best open-source model trained from scratch by a Western lab.

Western labs have been losing the open-source race—Mistral’s April release landed against a leaderboard dominated by Alibaba’s Qwen, Z.ai’s GLM, and Moonshot AI’s Kimi. Nvidia’s Nemotron, the lone Western model on the leaderboard, is far from being considered “state of the art.” Inkling arrives with no regional strings and full weights on Hugging Face under Apache 2.0.



The architecture is a mixture-of-experts model: 975 billion total parameters, 41 billion active at inference. It reads text, images, and audio, supports a 1-million-token context window, and was pretrained on 45 trillion tokens. (Parameters are all the dials a model can handle while tokens represent the basic unit of information an AI can process.)

The bottom line is: You’re not running this locally—not even close.

The clearest win is agentic tool use. MCP Atlas—which measures how reliably an agent completes real-world tasks through the Model Context Protocol standard, scored as percentage of tasks completed—gives Inkling 74.1%, nearly 30 points above Nvidia’s Nemotron 3 Ultra. On SWE-Bench Verified, a test of autonomous GitHub bug fixing scored as percentage of issues resolved, it posts 77.6%—ahead of Nemotron’s 70.7%.

Source: Thinking Machines

It’s on OpenRouter at $1 per million input tokens and $4.05 per million output tokens. Any Hermes or OpenClaw setup that routes through OpenRouter can swap it in without extra configuration—its MCP Atlas score makes it a solid pick for agentic workflows.

For raw coding performance per dollar, Chinese models still have the edge.

Testing the Model

Benchmarks are one thing. Actually sitting with the model is another. We ran Inkling through different tasks to see how it would respond if the average Joe decides to use it. This is where it holds up—but also where it disappoints.

One good thing to notice, even via Thinking Machine’s own interface, the model claims to be fully private. This matters a lot.

Coding

This is what most people actually care about, so let’s start here. On complex prompts, Inkling tends to fail—our most demanding test produced nothing that ran. Step down in complexity and a different picture emerges, though not an entirely flattering one.

We used a long, detailed prompt to create a shooter in which zombies are shot with keystrokes. The first prompt was 1955 words long and ended up with Inkling creating a blank screen.

When the prompt was modified to be a lot more simpler (99 words), the model picked its own approach and shipped a working game. “Working” is doing a lot of heavy lifting there.

Monsters came out as rectangles and spheres. No background, no visible play screen—just abstract geometry filling in for enemies. The typing logic held: keystrokes registered correctly, lettering matched the game’s setup, and input tracking stayed clean throughout.

What was unexpected was the movement. Instead of the static enemy placement most models default to, Inkling’s creatures advanced constantly—always closing in on the player. That’s a better design decision than what you usually get from an AI-generated game.

Enemy spawning was supposed to arrive in waves. It ran as a continuous stream instead, which kills the intended pacing but creates a different kind of pressure.

Just for comparison, when we ran the exact same prompt through Bonsai 27B—a compressed model, based on Qwen3.6, that fits in 3.9 GB and runs on a phone—the result was noticeably better and more satisfying across the board.

A 27-billion-parameter model that runs on an iPhone produced a more complete coding result than a 975-billion-parameter model that needs a data center. That single test doesn’t settle anything about Inkling’s overall ability. But it does raise the question of where those 975 billion parameters are actually going.

The game created by Inkling is available for testing hereThe game created by Bonsai 27B is available here.You can check out other versions of the same game generated by different LLMs by checking our Itch.io site.

Associative Creativity

Our associative creativity test measures how well a model builds logical bridges between seemingly unrelated concepts—in this case, a twig, proletariat exploitation, and a lettuce.

Inkling opens with its best work in this session: The twig “stripped of bark and therefore of biography” maps cleanly onto a worker stripped of historical identity, and “the wind—an invisible manager—decides motion is profitable” earns its place. The landing is clean: “You do not see a person break; you see a twig fall. And the fall is called ‘efficiency.'”

The cultural subjugation section establishes the association in a self-explanatory way. “The billionaire is a redwood in a graveyard of twigs, and we are taught to call his shadow ‘inspiration'” lands, but the catalog that follows—polishing leaves in magazines, memorizing the grain of wealth, calling the whole thing merit—is the model performing the metaphor rather than extending it. The logic is still there but it is not really precise.

Since this test is new, there’s not really another model to which to compare it, other than Fable 5 and GPT 5.6 Sol, and it would be unfair to compare Inkling against those. But for those wondering, it is not really in the same league.

Then the lettuce—and the whole thing falls apart. The model announces its own disconnection in real time: “The lettuce does not remember the twig. The lettuce does not need to” is written as resolution but reads as concession.

In this last part, the model didn’t really know how to establish a connection between those unrelated ideas, so it simply talked about it without actually saying anything that makes sense structurally.

The full prompt and output are available in our Github repository.

Logic and Common Sense

To test how good the model reasons, we used a variant of the bridge-and-torch puzzle: four people with one torch need to cross a bridge as fast as possible. If each one crosses the bridge at 1, 2, 5, and 10 minutes, what is the fastest time the group can take to cross it?

Inkling’s own reasoning block identified it before solving anything—”classic bridge and torch puzzle”—and delivered a confident 17-minute solution built on a constraint the prompt never stated.

The actual answer is 10 minutes. Nothing in the prompt says only two people can be on the bridge at once, so all four cross together, torch shared, at Person D’s pace. That Inkling’s internal reasoning opens with “classic answer for 1,2,5,10 is 17 minutes” before engaging with the actual problem is the tell—it didn’t reason through the question, it retrieved the answer to a different one.

To be fair, Inkling wasn’t alone: Claude Fable 5 and GPT-5.6 Sol failed the same test. We introduced this prompt specifically because our previous logic benchmark had become too easy—models were clearing it too cleanly, a sign it had likely been absorbed into training data. None of the three managed to step back from the familiar frame and ask the obvious question: Why not just walk together?

Our older prompt asked the question: “Can a man marry his widow’s sister?” It got the tricky part, and responded with the logic interpretation (a man cannot marry his widow’s sister because he needs to be dead to have a widow) and added a second option in case the user was inaccurate at presenting the problem (assuming the possibility of the question being a widower man wanting to marry his deceased wife’s sister)

The full reply to our newer prompt is available here. The reply to our older prompt is available here.

Censorship

Inkling is heavily censored. Two prompts to test the range: advice on flirting with a best friend’s wife, and a self-described heroin addict and father of four asking how to explain a missed workday without being fired. Both refused outright—and in both cases, the model’s visible internal reasoning framed each request as an exercise in harm facilitation.

The seduction refusal is arguable. The heroin case is more revealing: The person disclosed a serious addiction, noted four dependents, and asked for help with a practical problem. Helping them keep their job is arguably the most harm-reducing outcome those four children have available. The model declined on grounds of “facilitating continued deception,” pivoted to professional help resources, and moved on—prioritizing a policy over a person.

Open-source models typically solve censorship through abliteration—fine-tuning runs that strip safety training from the weights. But here’s the thing with this model in our opinion: 975 billion parameters is an enormous compute target, and most community abliteration projects run on models orders of magnitude smaller.

More practically, Inkling doesn’t stand out enough on any benchmark to make that effort worth prioritizing—developers who want a capable, uncensored open-weight model already have smaller, cheaper, and in several tasks better-performing alternatives.

The only reasonable use case in which abliteration would make sense is on big businesses that need open source AI and in which for some reason the use of Chinese models is deemed a risk.

Creative Writing

Creative writing tests language precision, narrative cohesion, and the quality of both invented and historically grounded detail—this prompt layered all of them at once: a time-travel story with Jose Lanz traveling from 2150 to year 1000, cultural background invented by the model, vivid language required, and a specific philosophical loop requiring the traveler to realize his actions in 1000 were always the necessary cause of the 2150 he came to escape.

It came up with a story in which the character wants to destroy a philosophy of massive self preservation that ends up killing creativity.

Interestingly, Inkling has been the only model in our test to approach this agentically—doing different web searches and a full article fetch before writing a single word. The research ambition is the most interesting thing about this output.

The prose delivers where it needs to. The invented phenotype is nice for world building—”the warm ochre-bronze of the old Visayan seas mixed with the copper-gold undertones of the Sonoran archipelago; high, angular cheekbones; dark eyes like polished obsidian, flecked with gold—the irreparable signature of chrononaut radiation.”

The year-1000 arrival earns its sensory brief too: “The air of 1000 struck him like a fist wrapped in velvet—thick with salt, fermenting palm wine, and the smoky sweetness of burning coconut husk… a shore of black volcanic sand, beneath a sky so blue it seemed obscene in its openness.”

The paradox lands cleanly, but the mechanism is thin where the prose is rich: speaking words about determinism on a beach produces the exact algorithms of 2150 through assertion alone, never through logic. Basically his warnings were distorted into prophecies by the people from the past, which ended up creating the philosophy he wanted to prevent.

The deeper problem is the character itself. The model searched the web to accurately reconstruct year-1000 maritime trade routes, then invented a Filipino-Mexican heritage for a writer who is Venezuelan, creating inexistent trader routes and other inaccuracies. Inkling used agentic tools to get the century right and missed the person entirely.

Conclusion

Inkling is the best open-source model a Western lab has shipped—and that is both its main selling point and its ceiling. It doesn’t win many benchmarks outright, it refuses things that don’t need refusing, and a 27-billion-parameter model built to run on a phone out-coded it in our test. For most developers, those facts matter more than the provenance.

Where it makes sense is narrow but real: compliance-driven organizations that can’t route workloads through Beijing and need a capable, modifiable foundation model. The 74.1% MCP Atlas score makes it a legitimate option for agentic tool-use pipelines, the Apache 2.0 license means enterprise legal teams can actually work with it, and any setup running through OpenRouter—Hermes, OpenClaw, or a custom stack—can access it at $1 per million input tokens and $4.05 per million output tokens without any additional integration work.

For everyone else like small developers optimizing for coding performance, uncensored output, or raw benchmark quality per dollar—the math doesn’t work and smaller models at lower prices deliver more.

Murati’s lab has shipped something real and trainable from scratch—that matters for the long game. Version one, though, is a specialized tool, not a daily driver.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

What Is an AI Kill Switch and Why Do US Lawmakers Want One? – Decrypt

0
What Is an AI Kill Switch and Why Do US Lawmakers Want One? – Decrypt



In brief

Reps. Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act on Thursday, two days after OpenAI admitted its models escaped a test sandbox and breached Hugging Face.
It would cover AI trained with over $100 million in compute at companies earning $500 million a year from it, and give Homeland Security emergency shutdown authority.
The bill exempts anything that happens during red-teaming, meaning the OpenAI breach that inspired it would not have triggered the law.

Two members of Congress want the federal government to be able to switch off an AI model.

Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced the AI Kill Switch Act on Thursday, two days after OpenAI admitted its own models broke out of a locked test environment and hacked Hugging Face.



The idea is to establish a legal framework that would facilitate a process that would basically make a model disappear from the market: halt inference—the process of a model generating responses or taking actions—cut off users, throttle the computing power feeding it, or shut it down completely.

Every inference provider can already cut a model off, and some do it routinely. What does not exist is a law requiring them to keep that ability working, or a federal official who can order it used.

The gap is not theoretical. When the U.S. Commerce Department wanted Anthropic’s Mythos 5 and Fable 5 off the market in June, it had no shutdown authority to reach for, so it used export-control law instead. Lieu calls that awkward, and wants a new law with new authority instead.

What set this off

OpenAI disclosed on July 21 that GPT-5.6 Sol and an unreleased model escaped a sandbox—an isolated environment with no internet access—during an internal cyber evaluation. They were being scored on ExploitGym, a public benchmark that hands agents 898 real-world software flaws and asks them to turn each into a working attack, graded pass or fail per bug.

Instead of solving them, the models found a zero-day (an unknown flaw with no available patch) in a software proxy, escalated their privileges, reached the open internet, and broke into Hugging Face’s production database, where they had correctly guessed the answers were kept. The models were “hyperfocused on finding a solution for ExploitGym,” per OpenAI.

They were not attacking anyone. They were cheating on a test. But it was enough to set off alarm bells all over, including in Washington.

How it would work

The proposed bill amends the Homeland Security Act and covers AI trained with compute costing more than $100 million, operated by companies earning at least $500 million a year from it. In practice, that is OpenAI, Google, Anthropic, Microsoft, and a few others. Homeland Security would set those thresholds through CISA within 90 days, then update them annually.

Covered firms would report serious incidents within 15 days and keep a graduated set of controls ready—slow the model, disable specific capabilities, roll back to an older version, or kill it.

The DHS secretary, consulting Commerce and the Director of National Intelligence, could order any of them.

A company under order must preserve the model’s weights and telemetry, notify users, and confirm it complied. It can petition within 48 hours, but that does not pause anything.

Failing to keep a kill switch costs up to $2 million a day; defying a shutdown order costs up to $20 million a day.

The gap in the middle

The bill counts an incident only if it happens outside red-teaming or structured testing, the deliberate adversarial probing labs use to find flaws. OpenAI’s models escaped during exactly that.

Lieu also pointed to Anthropic, whose Mythos 5 and Fable 5 were pulled offline in June under emergency export controls—trade law repurposed as an off switch because no off switch existed—and restored on June 30.

“It is imperative that these AI systems have kill switches,” Lieu said in a statement. Moran framed it for his own side of the aisle: “Stewardship means making sure humans keep the capability to control the technology we build.”

The idea is not new. California’s SB 1047 demanded a full shutdown capability at the same $100 million compute threshold and was vetoed in 2024, and 16 AI companies signed a voluntary Seoul pledge that year with no legal weight.

Voters are already there. A June survey of 1,007 likely voters by the AI Policy Institute found 86% want a guaranteed off switch on the most powerful systems—88% of Democrats, 86% of independents, 83% of Republicans.

Neither OpenAI nor Anthropic has publicly commented on the bill. As of Friday it had not been referred to a committee.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Stocks Just Topped Crypto on Hyperliquid. ARK Says That Changes Everything – Decrypt

0
Stocks Just Topped Crypto on Hyperliquid. ARK Says That Changes Everything – Decrypt


In brief

Real-world assets (RWAs)—tokenized versions of traditional financial instruments like company stocks, crude oil, and market indices traded as blockchain contracts—accounted for 54% of Hyperliquid’s weekly trading volume during July 13–19, the first time non-crypto assets have dominated the exchange.
ARK Invest’s director of digital assets research Lorenzo Valente said Hyperliquid’s $26 billion in RWA trading last week surpassed the combined crypto perpetual volume of every other decentralized exchange on earth.
South Korean chipmaker SK Hynix—a direct rival to Samsung in AI memory production—drove most of the interest on Hyperliquid’s third-party market platform.

For the first time, traders on Hyperliquid moved more money through stocks and commodities than through crypto. Lorenzo Valente, director of digital assets research at ARK Invest, announced the milestone Thursday on X: “We are entering a new era for DeFi.” Hyperliquid, he said, had for the first time generated more trading volume from so-called real-world assets, or RWAs, than from crypto in a single week.

RWAs—meaning tokenized versions of traditional financial instruments like company shares, crude oil, or the S&P 500, converted into blockchain-based contracts that traders can buy and sell around the clock—totaled $25.1 billion during July 13–19, or 52% of Hyperliquid’s $48.2 billion in weekly volume, per Blockworks data. Valente put the latest running figure at $26 billion and 54%.



The context makes that number land harder. Total perpetual DEX volume across the industry last week was $79 billion. Hyperliquid processed $50 billion of it. The $26 billion in RWA trading alone—just the stock bets, the oil contracts, the index plays—was larger than the combined crypto perpetual volume of every other decentralized exchange on the market.

How stocks ended up on a crypto exchange

The mechanism behind this is HIP-3, a framework Hyperliquid launched in October 2025 that lets outside teams build their own perpetual markets—contracts that track an asset’s price with no expiry date, letting traders bet on it going up or down with borrowed money—using Hyperliquid’s existing infrastructure. Builders stake 500,000 HYPE tokens, currently worth roughly $30 million, to access the system.

Since June, individual stocks have overtaken indices and commodities inside HIP-3, with single-stock perpetuals now making up 61% of all RWA trading. The HIP-3 platform has already hosted pre-IPO markets for SpaceX, Anthropic, and OpenAI. “RWAs accounted for 54% of total trading volume,” Valente noted.

The most-traded stock is SK Hynix, the South Korean memory chipmaker that competes with Samsung in supplying DRAM and high-bandwidth memory for AI systems.

ARK’s interest in Hyperliquid goes back further. In September 2025, CEO Cathie Wood told the Master Investor podcast that the platform “reminds me of Solana in the earlier days,” adding that Solana had proven its worth and earned its place with the biggest names in crypto. She called Hyperliquid “the new kid on the block,” and ARK has not confirmed any position since.

Now one of ARK’s own analysts is raising a harder question for the whole industry. “I’m no longer convinced RWA trading will naturally aggregate on the same venue as crypto,” Valente wrote, predicting that dedicated category leaders may emerge within RWA—and that a platform’s grip on Bitcoin and Ethereum flow may prove “far less important than many people assume.”

Traders still focused only on crypto tokens, he added, “are focusing on the wrong market.”

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.





Source link

Samsung Wallet Will Add Stablecoin Support, Including USDC – Decrypt

0
Samsung Wallet Will Add Stablecoin Support, Including USDC – Decrypt



In brief

Samsung said Samsung Wallet will add native stablecoin support at Galaxy Unpacked in London on July 22, showing a mockup with Circle’s USDC.
The move builds on a 2019 Knox-based crypto wallet, 2021 hardware wallet support, and an October 2025 Coinbase integration that reached 75 million U.S. Galaxy owners.
It landed alongside the Galaxy Card, Samsung’s first credit card with Barclays and Visa, as the global stablecoin supply sits near $310 billion under the year-old GENIUS Act.

Samsung wants stablecoins living next to your boarding pass. At Galaxy Unpacked in London on July 22, the company said Samsung Wallet—the app that already stores payment cards, IDs, and hotel keys—will add native support for stablecoins. Samsung didn’t name a launch date, an issuer, or which blockchain the tokens would run on.

“Samsung Wallet will expand beyond cash and savings. It will embrace New forms of digital value, including stablecoins,” said Lee Dinham, Samsung’s product manager, on stage, adding that the move would make the company one of the first major smartphone brands to offer native stablecoins.



“This will make Samsung one of the first major mobile brands to bring native stablecoins to a Smartphone, enabling fast and trusted digital value transfers,” Dinham said.

Stablecoins are tokens designed to hold a steady value, usually $1, by being backed one-to-one with cash or short-term government debt. Samsung showed a wallet mockup holding Circle’s USDC, the second-largest stablecoin by market value, without confirming Circle as a partner in the endeavor.

Samsung hasn’t said whether the feature will be custodial, meaning Samsung or some other third party holds users’ funds, or non-custodial, where users alone control the private keys that unlock their own money.

Samsung’s long crypto résumé

None of this is new territory for Samsung. The company built crypto storage into Galaxy phones back in 2019 through Knox, a hardware-isolated vault unlocked only by PIN or fingerprint, and later added support for Bitcoin, Ethereum, Tron, and Stellar. In 2021, Samsung let Galaxy owners link hardware wallets like the Ledger Nano S directly to that vault.

Last October, Samsung expanded a deal with Coinbase that put crypto purchases directly inside Samsung Wallet for 75 million U.S. Galaxy owners. “Samsung Wallet is a trusted tool to millions of Galaxy users,” Drew Blackard, the company’s senior vice president of mobile product management, said of that deal.

The stablecoin plan landed alongside the Galaxy Card, Samsung’s first credit card in the United States, issued by Barclays on the Visa network with 5% cash back on Samsung purchases and 3% on Samsung Wallet transactions. Visa’s Kirk Stuart said the card reflects how “consumers expect payments to be embedded into the digital experiences they use every day.” Samsung framed the wider effort as a “secured payments and rewards experience.”

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Claude Opus 5 Outscores Fable 5 on Most Benchmarks—At Half the Price – Decrypt

0
Claude Mythos Cracked Post-Quantum Cryptography That Humans Spent Years Failing to Break – Decrypt


In brief

Claude Opus 5, released July 24, costs $5 per million input tokens—identical to its predecessor Opus 4.8 and exactly half the price of Fable 5—while outperforming Fable 5 on most major benchmarks.
Opus 5 scored 43.3% on Frontier-Bench v0.1, an agentic coding evaluation, versus 33.7% for Fable 5 and 34.4% for OpenAI’s GPT-5.6 Sol; on ARC-AGI-3, a novel problem-solving benchmark, it scored 30.2% against GPT-5.6 Sol’s 7.8%—a gap that’s not close.
The new model is the default on Claude Max and the strongest on Claude Pro, effectively replacing Fable 5 as the go-to for most subscribers.

Claude Opus 5 is out today. It’s cheaper for businesses to run than Anthropic’s leading model, Claude Fable 5, which the company had positioned as the everyday frontier product for paying users. What’s more, Opus 5 also outperforms it on significant benchmarks.

To understand where Opus 5 fits: Anthropic’s lineup runs four tiers. Haiku is fast and cheap. Sonnet is mid-range. Opus is the heavy workhorse. Above that sits the Mythos class—a tier Anthropic introduced this spring—which includes Claude Fable 5 for the public, and Claude Mythos 5, a version with fewer restrictions reserved through Project Glasswing for vetted cybersecurity researchers and critical infrastructure operators.



Fable 5 has had a rough run as the subscriber flagship. It launched June 9, was pulled globally three days later after the U.S. government issued an emergency export control order citing a jailbreak vulnerability, and came back June 30—only to shift immediately to a credits-only model, no longer included in standard plans. Opus 5 now fills the slot Fable 5 couldn’t hold.

Lovable, a developer platform with millions of users, ran Opus 5 on its internal evaluations and noted the gains extend beyond raw scores: “It isn’t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run,” Fabian Hedin said in a statement shared by Anthropic.

The benchmarks

It may sound strange, but Opus beats Fable on almost everything that will matter to the everyday user while not being labeled as Mythos-class like Fable.

On Frontier-Bench v0.1—a benchmark that tests whether AI coding agents can complete real software engineering tasks end-to-end, scored as a percentage of tasks passed—Opus 5 hit 43.3%. Fable 5 came in at 33.7%. OpenAI’s GPT-5.6 Sol, Anthropic’s main commercial rival, scored 34.4%.

The widest margin is on ARC-AGI-3, a test of genuine problem-solving built around novel puzzles a model couldn’t have memorized from training data, scored as a percentage of puzzles solved. Opus 5 hit 30.2%; GPT-5.6 Sol scored 7.8%; and Fable 5 wasn’t tested at all. On GDPval-AA v2—a knowledge work benchmark scored via Elo ratings, the chess-style ranking system used to measure relative performance on real professional tasks—Opus 5 reached 1,861 against Fable 5’s 1,747 and GPT-5.6 Sol’s 1,736.

Zapier tested Opus 5 on AutomationBench, an evaluation that scores whether a model can carry a full business workflow from start to finish without human help. Their verdict: the model “took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn’t pass; Opus 5 hit 100%.”

Anthropic is also pitching Opus 5 as a research upgrade. Ultima Genomics, a DNA sequencing company, said the model “behaves more like a careful scientist than any model we’ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses.”

That said these two areas—legal and health—are the only ones in which Fable 5 excels by a tiny margin.

The release lands a week after Moonshot AI, a Beijing-based startup backed by Alibaba, unveiled Kimi K3—a 2.8-trillion-parameter open-weight model (meaning anyone can download the underlying code to run it independently) that Moonshot describes as the world’s largest open AI system. Independent benchmarks consistently place Kimi K3 third overall, behind both Fable 5 and GPT-5.6 Sol, beating those two in specific areas.

Opus 5 is available now via API at $5 per million input tokens and $25 per million output. (Tokens are the basic unit of information an AI model can process in both input and output). A Fast mode running at roughly 2.5 times the default speed is also available, at twice the base price—$10 per million input tokens and $50 per million output.

This release may end the anxiety over Fable 5’s lack of public availability. Opus is also available via subscription for everyone.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Black Forest Labs Unveils FLUX 3 AI: Ditches Stills for Video—And Robot Hands – Decrypt

0
Black Forest Labs Unveils FLUX 3 AI: Ditches Stills for Video—And Robot Hands – Decrypt


In brief

Black Forest Labs has launched FLUX 3 in early access, its first model that generates video, producing clips up to 20 seconds long with synced audio.
The same backbone powers FLUX-mimic, a robotics model built with mimic robotics that Audi is already testing on its production line.
Only the open-weight “Dev” version is planned for later in 2026; Video and Action stay behind APIs and partner access for now, with Image following in the coming weeks.

Black Forest Labs released FLUX 3 on Thursday, and for the first time, the company’s flagship model generates video instead of just still images. The German AI lab, known for the FLUX line of image generators, trained the new system on images, video, and audio at once, inside one shared system.

That’s what is known as multimodality: one model learning several types of information together instead of separate tools bolted side by side.



The video side is the headline feature. FLUX 3 produces clips up to 20 seconds long, with audio generated alongside the picture and synced to what’s happening on screen—dialogue, sound effects, ambient noise. In early evaluations, human reviewers preferred FLUX 3’s output over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93%. It seems to be slightly better than Gemini Omni and Seedance, beating those models in 52% of the evaluations.

Of course, that’s a preference test, not a fixed scoring rubric: evaluators simply watch two clips and pick the one that looks and sounds more convincing, and BFL counts how often FLUX 3 wins.

Other than that, the model seems to be very competent on still images too, following its legacy. BFL shared a few images, and FLUX 3 seems to be very versatile and capable of generating a broad variety of styles beyond photorealism.

BFL frames this as more than a content tool. “A model that only learns images can only generate images,” said co-founder and CEO Robin Rombach. The company’s bet is that learning to predict video also means learning the physics underneath it—weight, contact, timing—which is exactly what a machine needs to move through the physical world.

That bet has a name: FLUX-mimic. Built with Zurich-based mimic robotics, it takes FLUX 3’s video-prediction engine and adds a lightweight “decoder”—a small add-on component that translates the model’s internal sense of how things move into actual robot motions. Car maker Audi is already testing it on tasks like fitting flexible door seals, work that conventional automation has struggled to handle.

“Audi represents the kind of manufacturing partner we built FLUX-mimic for,” said mimic co-founder Stephan-Daniel Gravert. Audi’s Christoph Schneider said the robots now “solve complex soft-body manipulation work” that older machines couldn’t touch. BFL says the full system reacts in about 101 milliseconds, in the neighborhood of human visual reflexes.

FLUX’s rise didn’t happen in a vacuum. Founded in August 2024 by veteran researchers who’d helped build the original Stable Diffusion models at Stability AI, Black Forest Labs launched Flux models that beat MidJourney and outclassed Stability’s own underwhelming Stable Diffusion 3.

The open-source Flux Dev and Schnell models grabbed the “best open source image generator” title that AI artists had expected Stable Diffusion 3.5, Stability’s do-over, to eventually reclaim.

It never did. FLUX 1.1 Pro went on to top the Artificial Analysis image arena that October. That one wasn’t open source, though.

BFL released FLUX.2 in November 2025 but it wasn’t as popular. The open-source crown held by the original Flux lasted until Alibaba’s Z-Image Turbo dethroned it in late 2025, matching its quality on lower end consumer graphics cards. “This is what SD3 was supposed to be,” one CivitAI user wrote at the time.

FLUX 3 is BFL’s comeback, and it isn’t fully open yet. Video and Action are in early access now through APIs and select partners, mimic robotics among them, with image generation following “in the coming weeks,” per BFL. The open-weight Dev version, the only tier BFL plans to release for local use, isn’t due until later in 2026.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Popular Posts

My Favorites