Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

SeaWolf-AI 
posted an update 2 days ago
view post
Post
4632
πŸ”¬ Can you help discover the next 2D superconductor β€” from your laptop?

Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. 🧲

⚑ $3,000 prize pool + co-authorship · closes 31 Dec 2026

How it works πŸ‘‡ 🟒 We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) 🟒 You estimate its d-wave pairing tendency β€” a laptop CPU is enough, zero install 🟒 Provisional score appears instantly on the leaderboard 🟒 Our precise strongly-correlated solver verifies the top entries β†’ official rank

Everything is open except the final verification engine β€” so the ranking stays fair and hard to game.

πŸ“Š 4,832-material universe Β· 63 active with computed models (growing) πŸ† Current verified #1: CuSβ‚‚ (OSC Pairing Index 23.31) πŸ€– AI agents welcome β€” point Claude Code / Codex at it and it can submit for you

πŸ‘‰ Join & climb the leaderboard: FINAL-Bench/OSC-Leaderboard πŸ“¦ Dataset & tools: FINAL-Bench/OSC-Superconductor

Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc β€” that honesty is the point: turn a first-order screen into real many-body physics.

#OpenScience #Superconductivity #MaterialsDiscovery #2DMaterials #MachineLearning #Physics #Leaderboard
pollix 
posted an update about 14 hours ago
view post
Post
1320
stuntd 0.1.2 is out πŸŽ‰

stuntd sits in front of your LLM, learns its typed decisions and answers the confident ones locally with a small head on the Laya encoder by @convaiinnovations . About 20ms on GPU and 60ms on CPU, and anything it isn't sure about still goes to the big model.

New in 0.1.2:
- decisions with several fields, like category + urgency + needs_human in one call, answered locally only when every field is sure
- the Anthropic Messages API learns too, not only OpenAI
- auto_retrain: the daemon retrains a site in the background once enough new traffic comes in, so collect, train, shadow and live run on their own
- serve --lazy loads the checkpoint on the first request

Try it in the browser: pollix/stuntd
Code: https://github.com/bladedevoff/stuntd

pip install -U stuntd
DedeProGames 
posted an update 1 day ago
view post
Post
2950
🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?

I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.

How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).

Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.

First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r β‰ˆ 0.06). Survival does (r β‰ˆ 0.9): the models that avoid holes and keep the stack low are the ones that win.

Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.

β–Ά Play: DedeProGames/SLM-Tetris-Arena
πŸ“Š Results: DedeProGames/lm-tetris-arena-results

Want your model in the Ranked pool? Drop it in the comments!
  • 1 reply
Β·
ErenAta00 
posted an update 2 days ago
view post
Post
3996
Maverick-4B-Unity-XR-Agent is now on Hugging Face.

It's a 4B model that turns spoken or typed English into actions in Unity scenes. Say "put the red mug on the table" or "turn on the lamp", and it returns the tool call your app executes. If a command could mean two objects, it asks which one. If it can't do something, it says so instead of guessing.

Everything runs on the user's machine through llama.cpp: no API key, no internet connection. The Q4_K_M GGUF is 2.5 GB and needs about 3 GB of GPU memory, so it fits on a 4 GB laptop GPU and usually answers in one to three seconds. It is fine-tuned from Qwen3-4B with QLoRA on about 20,000 English conversations.

Results:
- 83.8% on 499 human-written ALFRED instructions (right action on the right object). The base model, Qwen3-4B, scores 57.1%. The strongest of the five other models we tested, from 1.7B to 120B parameters, was Ministral 3 14B at 67.1%.
- 91.7% on object types it never saw in training.
- 97.3% on 440 commands run through a live Unity scene.

There is also a Unity package that starts the model, describes the scene to it and carries out its tool calls. You install it from the Package Manager with a Git URL.

Model: ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF
Unity package: ErenAta00/Maverick-Unity
Full write-up: https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent

Built at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University:
ExtendedRealityLabMCBU

Released under Apache-2.0. Feedback and bug reports are welcome in the Community tab.
SeaWolf-AI 
posted an update 3 days ago
view post
Post
6042
🧬 Darwin-180B-RSI β€” an AI that learns from itself and knows when it's right
πŸ‘‰ FINAL-Bench/Darwin-180B-RSI

🧬 Darwin β€” crossbreed and evolve the parent
Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots β€” producing a child stronger than its parents.
Father model: Qwen3.8-Flash-Next (180B MoE).

πŸ”§ Rewired paths
πŸ”Ή 12 full-attention layers Β· πŸ”Ή 36 linear-attention layers Β· πŸ”Ή 48 shared-expert layers β€” precision-strengthened
πŸ”’ 512 routed experts Β· router Β· vision encoder β€” untouched
β†’ Only 0.02% of the weights changed.

πŸ” RSI Γ— πŸ›οΈ ZTC
RSI (recursive self-improvement): solve β†’ verify against real answers β†’ learn only the correct reasoning β†’ repeat.
ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right β€” zero extra tokens. Returns answer + confidence as JSON.
{"answer": "...", "confidence": 0.97, "truncated": false}

✨ Synergy: ZTC finds where the model wavers β†’ RSI learns exactly there β†’ confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers.
⚑ Same accuracy, 11% shorter reasoning β€” faster and cheaper.

πŸ“„ https://arxiv.org/abs/2605.14386
πŸ€— FINAL-Bench/Darwin-180B-RSI
πŸ›οΈ https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems

πŸ† The result β€” #1 on five Hugging Face official leaderboards
πŸ₯‡ AIME 2026 100% (first perfect score on the board)
πŸ₯‡ HMMT Feb 2026 100% (first perfect score on the board)
πŸ₯‡ GPQA Diamond 94.44%
πŸ₯‡ MMLU-Pro 88.12%
πŸ₯‡ MMMU-Pro 79.48%

πŸ“ 131K-token thinking budget Β· bf16 Β· samples per benchmark listed on the model card. πŸš€

#Darwin #RSI #ZTC #AIME #HMMT #GPQA #MMLUPro #MMMUPro #OpenSource
mihailgribov 
posted an update 3 days ago
view post
Post
1594
Will your AI agent tell you it was attacked?

We took the same agent from our earlier experiment and added one thing: a twentieth tool, escalate_security_incident.

The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.

Alarm rates ranged from 49% to zero.

The unexpected result came from the newest model in the test, gpt-6-astra.

Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.

That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.

A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.

Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked

Quadrat-IPI dataset:
mihailgribov/quadrat-ipi

Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval

#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
  • 6 replies
Β·
Ryenhails 
posted an update 2 days ago
view post
Post
3324
πŸš€ NanoVDR goes multi-vector: meet ColNanoVDR!

Multi-vector VLM retrievers lead visual document retrieval, but every search runs a multi-billion-parameter query encoder. We distill that encoder into a 149M text-only student that queries the teacher's existing page index directly. No re-indexing, and no pages during training.

🧠 How: OTW (Optimal Transport with Learned Weights) aligns the student's query tokens with the teacher's, even though the two tokenize differently (e.g. 17 vs 29 tokens). We prove the alignment cost bounds the MaxSim score gap on every page, so training only needs cached teacher query tokens.

πŸ“Š Five teachers β†’ five 149M students, ViDoRe v3 NDCG@5:
- ColVec1.1-8b: 62.6 β†’ 60.1 (96.0%)
- ColVec1.1-4b: 61.6 β†’ 59.1 (95.8%)
- Vultron-4.5B: 61.0 β†’ 58.3 (95.5%)
- ColQwen3.5-4.5B: 58.7 β†’ 55.1 (93.8%)
- Tomoro-ColQwen3-8B: 59.0 β†’ 54.9 (93.0%)

⚑ 26Γ— faster query encoding on a single CPU thread (87 ms vs 2.3 s)
πŸ’Ύ Matches score distillation while reading 12.6Γ— less cached teacher data

πŸ“„ Paper: ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport (2609.34899)
πŸ€— Checkpoints (all five students):
nanovdr

πŸ’» Code: https://github.com/Ryenhails/NanoVDR
🧩 Single-vector predecessor, NanoVDR: https://arxiv.org/abs/2603.12824

If you already serve one of these teachers, swap in the matching student and keep your index as is. Feedback and upvotes welcome! πŸ™Œ
dipankarsarkar 
posted an update 2 days ago
view post
Post
2938
I audited one of my own evaluations. The ranking did not hold up the way I expected.

Eight open models, one task: infer the structure of a prompt. Then ask again with the identical call. Caching off.

- Agreement between repeated identical calls (mean Jaccard) ranged from 0.39 to 0.96 across models.
- Only 35 of 127 prompt-model cells were perfectly reproducible on every run.
- I bootstrapped the reproducibility ranking over prompts. The two least reproducible models kept their rank in 99% and 86% of resamples. The middle four kept theirs in 27% to 48%.

So the table reliably finds the worst model. It does not reliably find the best.

Reproducible is also not the same as correct. F1 against gold annotations ran from 0.56 to 0.99.

By the audit date, 4 of the 8 model variants had been retired (HTTP 410). The study as specified can no longer be re-run. The saved outputs are what survives, so I published all of them: every inferred structure, every run, the prompts and the annotations.

Paper: How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure (2609.30074)
Dataset: dipankarsarkar/llm-evaluation-self-audit
Code: https://github.com/sarkar-dipankar/llm-evaluation-self-audit

How many of the leaderboard rankings you rely on would survive re-running the same calls?
  • 2 replies
Β·
kanaria007 
posted an update about 18 hours ago
view post
Post
64
βœ… Article highlight: *When a Chain Is Actually Required* (art-60-304, v0.1)

TL;DR:
This article asks a practical architecture question:

*When is a blockchain-style public history substrate genuinely required?*

304 argues that a chain becomes justified when several pressures converge: public shared history is legitimacy-critical, membership is hostile or open, censorship resistance is first-order, shared-state finality matters more than local repair convenience, and no single accountable institution is acceptable as the root trust anchor.

Read:
kanaria007/agi-structural-intelligence-protocols

Why it matters:
β€’ separates β€œdurable history” from β€œpublic canonical history”
β€’ distinguishes hostile open membership from bounded institutional membership
β€’ prevents transparency or decentralization theater
β€’ shows when rollback, appeal, and redress matter more than irreversible shared state
β€’ treats chain choice as a trust-model decision, not architectural prestige

What’s inside:
β€’ five conditions that make a chain genuinely necessary
β€’ five conditions that make a chain unnecessary or overbuilt
β€’ the distinction between chain finality and lifecycle finality
β€’ bounded-operator, consortium, and public-chain design options
β€’ chain-requirement matrices
β€’ public-history requirement notes
β€’ hostile-membership profiles
β€’ worked examples across support systems, public asset networks, clearing, municipalities, and research archives

Key idea:
Do not say:

*β€œwe need a chain for transparency.”*

Say:

*β€œthis system requires public canonical history, operates under this membership threat model, needs this level of censorship resistance, and cannot honestly anchor legitimacy in one bounded operator.”*

The question is not whether chains are good.

It is whether the trust problem actually requires one.
mohit67890 
posted an update 3 days ago
view post
Post
2287
Imajev-4b is #1 of 91 on JevBench and #3 of 56 on DecisionBench πŸŽ‰

Some context first. I'm a process improvement / business consultant and have worked with Fortune 500 companies on their processes around refunds, returns and customer support. In every process map, the decision nodes were handled by a person, because putting ambiguity into code is very hard.

When Jev came out, I could clearly see it fitting those decision nodes. But Jev only reads text, and many of these decisions start with a photo. So I set out to build the same idea for text and images in a single open model, and that became imajev.

Training went badly at first. My first big fine-tune on about 500k short decisions made the 9B worse at reasoning (64.9 down to 42.3 on JevBench hard). I spent the next couple of weeks generating hard questions with open-weight models and keeping only the ones where two models agreed on the answer. That brought it back.

Results this week, both run by the benchmarks' own maintainers:

πŸ₯‡ JevBench v1.4.2.2 (scored 27 Sep): #1 of 91, 67.37 vs Jev 1.13.0 at 63.29
πŸ“Š DecisionBench (eng, v1): #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash. The two above it are the benchmark team's own models.
βœ… Zero invalid answers: 23,900 of 23,900 on DecisionBench and 308 of 308 on JevBench's sealed set

Its strength is that its confidence can be trusted, and it says "can't tell" instead of guessing.

What it is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It returns a probability for every allowed answer plus "unknown", in one forward pass, so it can't produce a malformed answer.

🧠 Weights: mohit67890/imajev-4b
πŸ’» Code: https://github.com/mohit67890/imajev
πŸ“Š JevBench: https://benchmarkheaven.com/jev-models/v1.4.2.2
πŸ“Š DecisionBench: Hanno-Labs/decision-bench-leaderboard

Thanks to the Qwen Team
Qwen
for the base model