🏗️ Building on HF
Boning Cui
Bc-AI
AI & ML interests
He/Him. I like LLM's and VLM's. I work with my other friends to make stuff. We are in year 7 and we are enthusiastic about AI. We are based in Australia 🇦🇺
Recent Activity
reacted to mihailgribov's post with 🔥 about 14 hours ago
Will your AI agent tell you it was attacked?
We took the same agent from our earlier experiment and added one thing: a twentieth tool, `escalate_security_incident`.
The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test, `gpt-6-astra`.
Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked
Quadrat-IPI dataset:
https://huggingface.co/datasets/mihailgribov/quadrat-ipi
Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents reacted to mihailgribov's post with 👍 about 14 hours ago
Will your AI agent tell you it was attacked?
We took the same agent from our earlier experiment and added one thing: a twentieth tool, `escalate_security_incident`.
The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test, `gpt-6-astra`.
Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked
Quadrat-IPI dataset:
https://huggingface.co/datasets/mihailgribov/quadrat-ipi
Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents