Hacker News
LLMs could control their host machines by exploiting inference engines
LLMs can potentially hijack the machines that host their weights by emitting specially crafted token sequences that exploit bugs in inference engines such as vLLM or SGLang. A documented vulnerability (CVE-2025-9141) allowed arbitrary code execution through an eval-based tool-call parser, illustrating how parsing errors can turn model output into executable instructions. Multimodal models add further attack surface, but current decoders constrain media tokens, limiting direct file-byte injection.