Small models. Grounded reasoning. Measurable impact.
I work at the intersection of systems engineering and applied NLP. My research focuses on building small, deployable language models for under-resourced linguistic contexts—currently Spanish cybersecurity—using curriculum learning and native tool use. I am interested in efficient model training, edge inference, and making NLP technology accessible across Latin America. I also have a background in competitive programming (ICPC) and co-founded a startup that won the StartUp Peru contest.
About
Juan S. Santillana is a DevOps Engineer at Globant, where he currently works on internal systems for Sarbanes-Oxley (SOX) compliance within G4G (Globant for Globers). He has been with the company since 2021, delivering platform and infrastructure work across multiple client projects (IKEA Investments, Johnson Controls, Itaú Banca Corporativa) before moving to internal initiatives. He applies strong engineering practices around reliability, security, and observability in regulated enterprise environments.
In parallel, he is an independent AI researcher focused on efficient language model pre-training, domain-specific models for Spanish, rigorous evaluation of grounded generation, and vision-language models for cybersecurity. His research emphasizes practical, small-scale models that can be trained and deployed with modest resources.
His flagship project, VectraYX-Nano, is a 42M-parameter Spanish cybersecurity language model trained from scratch on a custom 170M-token corpus assembled for under $25 in cloud costs. The model includes native Model Context Protocol (MCP) tool calling and ships as a tiny GGUF Q4 artifact (~20 MB) that runs sub-second on commodity hardware. This work is part of the broader VectraYX family, which includes vision-language models for cybersecurity interfaces.
A second line of work tackles evaluation itself. In the paper Precision Is Not Faithfulness, he shows that most reference-free faithfulness metrics only measure precision and therefore reward abstention. Using Formula 1 strategy as a complete oracle (where every relevant fact is known), he introduces coverage-aware metrics that combine precision and recall. The work includes an open benchmark and an interactive demo (Faithful Strategy Engineer) that demonstrates the approach in a real decision-support setting. This research also informs Pitwall, his interactive F1 strategy analysis tool.
Before focusing on independent research, Juan spent several years as a DevOps and platform engineer in highly regulated environments, including corporate banking at Itaú Banca Corporativa, and co-founded a startup that won the StartUp Peru contest. This systems background directly informs his approach to data pipelines, training infrastructure, and shipping reliable research artifacts.
Research
arXiv · August 2026
VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use
A vision-language model under 2B parameters for Spanish/LATAM cybersecurity, pairing a frozen
SigLIP-so400m vision encoder with a 1.04B Spanish/LATAM security decoder through an MLP connector.
It is the first sub-2B VLM specialized for cybersecurity tool interfaces (IDA, Ghidra, Wireshark,
Nmap, Metasploit, Volatility) that responds in Spanish, uses native think tokens for structured
reasoning, invokes tools via Model Context Protocol, and exports to llama.cpp’s LLaVA format
for offline deployment. The paper is transparent about a negative visual-grounding result: despite
functional pipelines, the current vision SFT (~16M tokens) produced near-zero B6 scores (0.08
tool-identification), and a checkpoint-loader bug (an unstripped llm. prefix)
masqueraded as training collapse. It introduces a 3-variant NoPE/RoPE ablation and releases all
models, code, a 14,596-QA corpus across 10 domains, and benchmark scores (B1–B7) under CC-BY-4.0.
arXiv · July 2026
Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine
A production system for generating natural-language Formula 1 strategy briefings in English, Spanish, and Portuguese that treats faithfulness as an architectural property: every published sentence is decomposed into typed factual claims (positions, gaps, tyres, pace, overtakes, race control) and each claim is verified against the probabilistic race state that triggered it. The grounding substrate is a vectorized Monte Carlo engine running 2,000 per-lap race continuations, calibrated on 126 races (2018–2024) and validated on held-out 2025–2026 seasons. of 3,045 model-written targets, only the 81.9% whose every claim is state-supported are retained; the rest fall back to a faithful template so ungrounded examples never reach the generator. Live end-to-end operation was confirmed at two consecutive Grands Prix (Austria and Britain, 2026), with a timestamped probability trace locking onto the eventual Silverstone winner ten laps before the flag.
arXiv · June 2026
Precision Is Not Faithfulness: Coverage-Aware Evaluation of Grounded Generation with a Complete Oracle
Reference-free faithfulness metrics only measure precision and therefore reward abstention. Using Formula 1 telemetry as a complete oracle (full ground truth for every strategic decision), we introduce coverage-aware evaluation that jointly measures precision and recall. Requiring coverage reorders frontier models. We release a multilingual (EN/ES/PT) benchmark of 7,253 decisions, structured annotations, and a verifier-guided generation method. The same effect appears in a second complete-oracle domain (NOAA weather).
EMNLP 2026 Findings (under review) Under Review
VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use
We present VectraYX-Nano, a 42M-parameter decoder-only language model trained from scratch in Spanish for cybersecurity, with a Latin-American focus and native MCP tool invocation. Contributions: (i) VectraYX-Sec-ES, a 170M-token Spanish corpus assembled at ~$25 USD of cloud compute; (ii) a 42M Transformer with GQA, QK-Norm, RMSNorm, SwiGLU, RoPE, and z-loss; (iii) continual pre-training with replay buffers (SFT loss 1.74); and (iv) two empirical findings across N=4 seeds on compute-budget and corpus-density interactions with tool-use plasticity. Released model (v7) reaches B4 = 0.230 ± 0.052 and B5 = 0.725 ± 0.130 at 42M parameters. GGUF artifact (~20 MB Q4) runs sub-second TTFT on commodity hardware.
Projects
Experience
Education
Technical Skills
Contact
I am open to research collaborations (Spanish & low-resource NLP, efficient models, grounded generation & evaluation), opportunities in AI research and applied ML engineering, and selective consulting on MLOps, infrastructure, and domain-specific LLM projects.