COLIDE
Custom CUDA kernels for a CNN-BiLSTM intrusion-detection model, with an on-device LLM that explains alerts without ever sitting in the detection path. A systems study — and an exercise in not overclaiming.
- CUDA
- HPC
- IoT Security
- LLM
- Systems
- BoT-IoT macro-F1
- 0.9780 ± 0.0033
- Sealed multi-seed test, n=5 (seeds 42–46)
- Blocks 1/2/4 vs matched PyTorch
- 3.24×–6.55×
- Operator-for-operator, RTX 3050
- Block 3 kernel progression
- 7.55×–9.50×
- Naive → FP16 half2 gate packing, five sessions
- Alert dispatch
- 16.60 µs p99
- Dispatch only — generation runs off the detection path