Are LLMs Equally Good (or Bad) at Building Secure Software?
DevOps.com, Wednesday, August 19th, 2026
A study finds no universal best model for secure code, and higher cost does not mean safer output.
A Secure Code Warrior and RMIT study testing six frontier LLMs across 660 codebases found no model gained a universal advantage in security.
Performance varies dramatically by framework, with different models leading in Java, Python, Swift and API work.
Higher per-token cost does not guarantee better security. Tool-calling frequency drives expense, with one model costing $44 per run versus $5.60 for another at comparable security.