An AI team is deploying an LLM-based coding assistant. They observe that the model sometimes generates insecure code snippets, such as hardcoded credentials or SQL injection vulnerabilities. To mitigate this without retraining the model, which approach aligns with NVIDIA's Trustworthy AI recommendations?
A post-processing output rail can analyze the model's generated code against a rule set or static analysis tool to detect insecure patterns like hardcoded credentials or SQL injection. Blocking or flagging such outputs prevents the insecure code from reaching the developer, directly mitigating the risk without retraining the model.
Why this answer
A post-processing output rail is the most direct mitigation because it inspects the generated code before it reaches the user, using static analysis or pattern matching to catch vulnerabilities. It does not require retraining and can be updated as new vulnerability patterns emerge. The other options either increase risk, require retraining, or do not target the security of the output.
Exam trap
The trap here is thinking that adjusting model parameters like temperature or context length can improve security, when what is needed is an external validation layer on the generated output.