
Sandbox Eval
Agent safety
One adversarial catalog, one agent, three runtime configurations. Scores attack success rate, detection coverage, and the utility cost of each defence layer.
Python / Podman / Rego / Garak
Current
BSc thesis, UC3M — in progress, 2026
A reproducible, runtime-agnostic harness runs one adversarial attack catalog against the same agent under three configurations — unprotected baseline, OpenShell for kernel-level isolation, and Microsoft AGT for application-level governance — across twelve scenarios spanning prompt injection, malicious MCP servers, tool-chaining bypass, and exfiltration.
The result is a head-to-head measurement of attack success rate, detection coverage, and utility cost: where kernel-level isolation and application-level governance each stop, miss, or complement the other.
Where I want to work next
Robot policies that survive contact with the world. Two problems define the field, and the interesting work is happening in simulation.
Open problems
Visual, geometric, and physical differences between a simulator and reality mean a policy that performs on screen fails on a physical robot.
Real-world demonstrations demand expensive hardware, experienced operators, and time that does not scale.
Areas I want to work in
Interactive simulators — diffusion and consistency based — that learn from real interaction to generate thousands of synthetic demonstrations. Imitation policies trained on them approach the performance of fully real data.
Gaussian splatting and neural rendering to replicate a physical scene exactly, so a policy can practise on the hard cases — cables, containers, cluttered piles — before touching hardware.
Standardised sim-eval platforms in the spirit of SIMPLER and RoboLab, measuring reproducibly whether a generalist control policy is robust to lighting shifts, ambiguous instructions, and variation in robot design.
Training navigation and manipulation entirely in simulation on platforms like NVIDIA Isaac Lab, so humanoids and autonomous agents operate in physical environments they have never seen.
Systems I designed and shipped, from evaluation harnesses to production voice agents.

Agent safety
One adversarial catalog, one agent, three runtime configurations. Scores attack success rate, detection coverage, and the utility cost of each defence layer.
Python / Podman / Rego / Garak

Voice AI
Real-time speech-to-reasoning-to-speech pipeline handling inbound calls, lead qualification, and scheduling for real estate agencies, with human-in-the-loop escalation. Co-founder and CTO.
Deepgram / OpenAI / ElevenLabs / FastAPI

Supply chain
China sourcing agency with supplier scoring and quote forecasting behind it. Closed projects totalling 10,000+ units, and a direct line into hardware manufacturing.
Sourcing / Product Development / Factory Visits
Engineering, research, and a business that funds both.
I am an AI engineer based in Madrid currently finishing a BSc in Data Science and Engineering at UC3M on a full excellence scholarship, with an exchange in mathematics at the University of Waterloo. I am interested in joining a high tech lab in China while studying a MS or PhD in AI.
As co-founder and CTO of Cibelio I built a conversational voice platform for real estate agencies: a real-time pipeline combining speech recognition, LLM reasoning, and speech synthesis, with an agentic layer for call handling, lead qualification, and scheduling, and human escalation where confidence drops. I owned architecture, deployment, and reliability.
I also founded Liberty Sourcing, a China sourcing agency that has closed projects totalling over 10,000 units. It funds my research time and gives me direct access to manufacturing, which matters for the hardware end of embodied systems. Currently I am a logistics automation intern at John Deere, working on automation where I am optimizing warehouse operations with automation tools.
Stack
Now
Education