A scanner can tell you an application has an injection flaw. Nothing in the standard toolchain can tell you whether a language model will do something it was instructed not to do, because the answer is not a property of the code. It is a property of the model, the prompt, the surrounding system, and how hard someone is willing to push.
That gap is the reason for this work: assessment tooling aimed specifically at applications with a language model behind them, built to probe the behavior rather than the source.
It is in development and deliberately not linked here yet. When there is something worth showing, it will be on this page.