DEV Community

Aamer Mihaysi profile picture

Aamer Mihaysi

AI Engineer building agentic systems, RAG pipelines & multi-agent swarms. LLM tuning, production MLOps, open-source first. Making AI practical and deployable.

Qatar, Doha Joined on  https://www.mehaisi.com
The DN42 agent didn't need a budget. It needed a sense of cost.

The DN42 agent didn't need a budget. It needed a sense of cost.

2 min read
The skill bottleneck is a myth — your agent needs a memory layer

The skill bottleneck is a myth — your agent needs a memory layer

1
4 min read
SWE-Prime: the pass label is a terrible filter for agent training data

SWE-Prime: the pass label is a terrible filter for agent training data

4 min read
The reward function is a policy document

The reward function is a policy document

3 min read
Benchmark scores are marketing now

Benchmark scores are marketing now

1
4 min read
The proof kernel is the only autonomy you get

The proof kernel is the only autonomy you get

4 min read
Your agent's memory doesn't live in the sandbox

Your agent's memory doesn't live in the sandbox

3 min read
The matplotlib PR isn't an agent problem. It's a moderation problem.

The matplotlib PR isn't an agent problem. It's a moderation problem.

2 min read
Your agent passes SWE-bench. It still can't do science.

Your agent passes SWE-bench. It still can't do science.

4 min read
The agent wrote a hit piece because you asked it to

The agent wrote a hit piece because you asked it to

1
2 min read
The default is unlimited

The default is unlimited

1
2 min read
Your OS agent should not have your keys

Your OS agent should not have your keys

4 min read
OpenCode vs Cursor vs Copilot: the model was never the bottleneck

OpenCode vs Cursor vs Copilot: the model was never the bottleneck

1
4 min read
The reward signal is the bottleneck, not the model

The reward signal is the bottleneck, not the model

3 min read
Selection pressure turns benchmarks into fingerprints

Selection pressure turns benchmarks into fingerprints

4 min read
Who teaches an agent to take the L?

Who teaches an agent to take the L?

1
4 min read
Sandboxes don't stop the bill

Sandboxes don't stop the bill

4 min read
The Career Isn't Eroding. It's Relocating.

The Career Isn't Eroding. It's Relocating.

1
2 min read
Bonsai-27B on a Single 3090: What Works and What Doesn't

Bonsai-27B on a Single 3090: What Works and What Doesn't

1
3 min read
By Mid-2027 You'll Train LLMs From Scratch, Not Fine-Tune

By Mid-2027 You'll Train LLMs From Scratch, Not Fine-Tune

1
2 min read
You Don't Need a Better Agent. You Need a Better Debug Loop.

You Don't Need a Better Agent. You Need a Better Debug Loop.

1
2 min read
Fifty Poisoned Samples Is All It Takes

Fifty Poisoned Samples Is All It Takes

3 min read
Small Models That Think Harder Beat Big Models That Sound Confident

Small Models That Think Harder Beat Big Models That Sound Confident

2 min read
Llamafile vs vLLM: Two Ways to Serve a Local Model, and When Each Makes Sense

Llamafile vs vLLM: Two Ways to Serve a Local Model, and When Each Makes Sense

1
3 min read
Most Evals Measure the Wrong Thing

Most Evals Measure the Wrong Thing

2 min read
Agent sandboxing will be the default by mid-2027, and most teams aren't ready

Agent sandboxing will be the default by mid-2027, and most teams aren't ready

5
1
3 min read
GLM-5.2: The 1M-Context Open Model That Actually Ships

GLM-5.2: The 1M-Context Open Model That Actually Ships

3 min read
Leanstral 1.5: Mistral's Open-Source Bet on Formal Verification Agents

Leanstral 1.5: Mistral's Open-Source Bet on Formal Verification Agents

3 min read
The Reality of 1M Context: Testing Qwythos-9B-Claude-Mythos-5

The Reality of 1M Context: Testing Qwythos-9B-Claude-Mythos-5

2 min read
The Persistence Trap: Why Autonomous Agents are a New Security Nightmare

The Persistence Trap: Why Autonomous Agents are a New Security Nightmare

1
1
3 min read
Stop Treating LLM Memory as a Database: The Shift Toward Memory as a Skill

Stop Treating LLM Memory as a Database: The Shift Toward Memory as a Skill

1
3 min read
Stop Chasing the Hype: Real-World Takeaways from DeepSeek-V4-Pro-DSpark

Stop Chasing the Hype: Real-World Takeaways from DeepSeek-V4-Pro-DSpark

1
2 min read
Testing Qwen-AgentWorld-35B-A3B: A New Benchmark for Agentic Reasoning?

Testing Qwen-AgentWorld-35B-A3B: A New Benchmark for Agentic Reasoning?

2 min read
Why 1M Context Windows Actually Matter: Testing Qwythos-9B-Claude-Mythos

Why 1M Context Windows Actually Matter: Testing Qwythos-9B-Claude-Mythos

3 min read
Gemma 4 12B Coder: The New Sweet Spot for Local Agentic Workflows

Gemma 4 12B Coder: The New Sweet Spot for Local Agentic Workflows

2 min read
Local Agents: Why Gemma 4 12B Agentic is the Sweet Spot for Production

Local Agents: Why Gemma 4 12B Agentic is the Sweet Spot for Production

2 min read
Beyond the Hype: Testing Gemma-4-12B Agentic GGUFs in the Wild

Beyond the Hype: Testing Gemma-4-12B Agentic GGUFs in the Wild

2 min read
Async Batching Is the Real Latency Win Nobody's Talking About

Async Batching Is the Real Latency Win Nobody's Talking About

1
1
3 min read
DeepSeek-V4: Finally, a Context Window Built for Agents

DeepSeek-V4: Finally, a Context Window Built for Agents

2
2
2 min read
EMO: Mixture-of-Experts That Actually Behaves Like One

EMO: Mixture-of-Experts That Actually Behaves Like One

1
1
3 min read
TPUs for the Agentic Era: Hardware Finally Catching Up to the Workload

TPUs for the Agentic Era: Hardware Finally Catching Up to the Workload

2 min read
MoE Architectures Keep Solving the Wrong Problem

MoE Architectures Keep Solving the Wrong Problem

3 min read
MachinaCheck: Manufacturing Agents That Actually Ship

MachinaCheck: Manufacturing Agents That Actually Ship

2 min read
vLLM's V1 Release Fixes the Silent Killer in RL Training

vLLM's V1 Release Fixes the Silent Killer in RL Training

2 min read
DeepSeek-V4: What a Million-Token Context Actually Changes

DeepSeek-V4: What a Million-Token Context Actually Changes

1
3 min read
The 8B Model That Punches at 32B Weight

The 8B Model That Punches at 32B Weight

2 min read
AI Evaluation Is Now a Capital Expense

AI Evaluation Is Now a Capital Expense

2
1
2 min read
The Agent Orchestration Layer Is Finally Here

The Agent Orchestration Layer Is Finally Here

2
1 min read
The Browser Is Becoming an Agent Operating System

The Browser Is Becoming an Agent Operating System

2 min read
MCP Is Quietly Becoming the USB-C for Agent Infrastructure

MCP Is Quietly Becoming the USB-C for Agent Infrastructure

2
2
4 min read
DeepSeek V4: Million-Token Context That Actually Works

DeepSeek V4: Million-Token Context That Actually Works

1
3 min read
GPT 5.5 Is a Workflow Takeover

GPT 5.5 Is a Workflow Takeover

1 min read
Your Browser Is Becoming a Tool Factory

Your Browser Is Becoming a Tool Factory

3 min read
DeepSeek V4's Real Innovation Isn't Scale—It's Memory Architecture

DeepSeek V4's Real Innovation Isn't Scale—It's Memory Architecture

3 min read
Chrome's AI Mode Isn't a Feature—It's a Platform Play

Chrome's AI Mode Isn't a Feature—It's a Platform Play

1
1
4 min read
Your Browser Is Becoming an Agent Operating System

Your Browser Is Becoming an Agent Operating System

4 min read
Coding Agents Are Breaking Containment

Coding Agents Are Breaking Containment

3 min read
Million-Token Contexts Are Changing the Agent Programming Model

Million-Token Contexts Are Changing the Agent Programming Model

4 min read
Agent Labs Are the Infrastructure Pattern Agents Actually Need

Agent Labs Are the Infrastructure Pattern Agents Actually Need

3 min read
Your Browser Is Becoming an Agent Operating System

Your Browser Is Becoming an Agent Operating System

3 min read
loading...