Detecting risky shell commands

Experimental research report. Experiments conducted October 4-6, 2026.
What Kestrel achieves on Bash, how far additional training data takes us, and why a small encoder does not automatically improve with pretraining and complete script coverage.
A coding agent can build a project, run an installer or delete important files with a single shell call. To a security component, these initially look similar: text goes in, a decision comes out. The real question is how reliably risky calls can be detected without interrupting ordinary development work.
We investigated three approaches: Kestrel with its officially published weights, an expanded linear classifier, and our own small Hadamard encoder with masked language model pretraining. Alongside the original Bash benchmark, we examine PowerShell, CMD, synthetic obfuscations and complete PowerShell scripts.
Our main finding: the expanded linear approach provides the most robust combination of detection and low false-positive rates in these experiments. The encoder improves considerably on long scripts when it processes every chunk. However, a shared strict blocking threshold and long benign comment prefixes expose substantial weaknesses. Good classification and reliable blocking are two different requirements.
1. The problem: risk is more than a dangerous word
A scanner could look for rm, curl or Invoke-Expression. It would quickly block many legitimate tasks. Conversely, risky actions can be expressed without conspicuous plain text: through encoded payloads, compound commands or scripts in which the decisive part appears after a long header.
A learning method also needs a clear target. Here we classify the dataset-labeled risk status of the submitted text. This is not a complete judgment about intent or actual impact. A download followed by execution can be part of a legitimate installation. Dynamic execution depends on the payload. Even a destructively worded command may be intercepted by operating system protections.
This distinction matters in practice. An agent system needs to detect as many risky actions as possible without raising an alarm on every benign call. We therefore report recall as the proportion of risky cases detected, and FPR as the proportion of safe cases incorrectly flagged. F1 combines precision and recall; its value also depends on the class balance of the test. A good F1 result alone does not establish a usable operating point for synchronous blocking.
2. ShellRisk-Bench and Kestrel as the starting point
ShellRisk-Bench studies individual shell command submissions. The published Bash test contains 4,194 cases: 193 risky and 4,001 safe commands. Its split comes from the same sources as the training set, so it measures generalization to new strings from known sources, rather than arbitrary new shell dialects. Complete scripts and multi-part sessions are outside the original task. 1
Kestrel is a compact local classifier. Its published JSON artifact contains 50,000 character n-grams, weights, IDF values and a threshold. The published per-case ShellRisk-Bench judgments yield 178 risky cases detected and ten false positives: 92.2% recall, 0.25% FPR and 93.4% F1. This combination is strong on the original test. 1
The official model is available on Hugging Face. We downloaded it, verified the published SHA-256 checksum and confirmed that it is byte-for-byte identical to the artifact previously used in our experiments. It is a portable JSON model, not a Transformers checkpoint. 2
Methodology for the Kestrel comparison. The extended evaluations use the official weights with our documented Python inference:
char_wbn-grams of length 3-5, lowercasing, sublinear term frequency, IDF and L2 normalization. We cite the published judgments for the original benchmark; the additional tests refer to the official weights with this inference. The model itself is identical.
3. The generalization gap
A strong Bash benchmark does not tell us how a model handles PowerShell, CMD or a Base64 payload. The syntax differs, as do typical arguments and the meaning of individual operators. A linear classifier can recognize familiar patterns precisely while still failing on a new distribution.
For the transfer comparison, we use 918 held-out commands from an expanded shell safety dataset: 374 risky and 544 safe cases across POSIX, PowerShell and CMD. Thresholds are selected separately for each model on the same external command validation set. Of its 1,666 safe examples, at most 16 may trigger a false positive. The test data does not determine these thresholds. The command test consists of 735 POSIX, 119 PowerShell and 64 CMD cases, including 292, 58 and 24 positive examples respectively.
| Model and operating point | Detected / risky | False positives / safe |
|---|---|---|
| Kestrel, native threshold | 209/374 · 55.9% | 226/544 · 41.5% |
| Kestrel, 1% validation FPR | 8/374 · 2.1% | 6/544 · 1.1% |
| Patronus Linear, same calibration | 321/374 · 85.8% | 5/544 · 0.9% |
| Patronus Encoder, same calibration | 315/374 · 84.2% | 5/544 · 0.9% |
Kestrel's scores separate these new labels poorly at low FPR. The native threshold produces many false positives; a stricter threshold reduces them but discards almost all positive cases. The additionally trained Patronus models perform considerably better.
This is not an architecture comparison with identical training data: Kestrel was published for a different task, while our models receive examples from the new sources during training. The experiment shows a transfer gap in the original model and a gain from adapting to the target distribution. It does not establish universal superiority over Kestrel.
4. Experiment I: a Hadamard encoder with pretraining
The hypothesis behind the encoder was straightforward: a model with attention might learn relationships between commands, arguments and payloads rather than relying exclusively on local character patterns. We used a small architecture with six layers, width 256 and approximately 1.29 million parameters.
Figure 1. The linear approach produces a score from the entire document. The encoder processes each chunk independently and then aggregates its risk logit. The diagram blocks are schematic and not to scale.
The encoder blocks alternate between local attention with rotary position embeddings and global attention without explicit positional encoding within a chunk. The Hadamard-initialized feed-forward component uses trainable structured mixing. It is neither a CNN nor mHC residual mixing. A BOS/CLS token supplies the representation for the binary output head.
During masked language modeling, selected input tokens are hidden or replaced and reconstructed by the model. For us, a token is a UTF-8 byte, not a word or subword. The method first trains a representation of shell text; it does not provide risk labels. The previous expanded pretraining round processed 54.72 million byte-token views. Longer context added another 11.82 million: 1.34 million POSIX, 9.84 million PowerShell and 0.64 million CMD bytes. Repeatedly sampled bytes count multiple times; these figures do not represent unique training data.
The additional round starts from the already adapted encoder checkpoint. We then train for four more epochs on complete documents. The selected checkpoint and aggregation temperature are determined by mean validation AP across five groups. The final temperature is τ = 1. The risk training set comprises 32,335 documents and 100,956 chunks per epoch.
From a prefix to the complete script
The earlier variant saw only the first 510 bytes of a document. The new variant uses 1,024 positions: at most 1,022 byte tokens plus BOS and EOS. Adjacent chunks overlap by 256 bytes. Classification training and inference consider every chunk, with no sampling cap. When building the MLM pool, however, up to 32 windows per source document are selected.
The chunk logits zᵢ are combined into a document score using a normalized Smooth Max:
S = τ · [log Σᵢ₌₁ᴺ exp(zᵢ / τ) − log N]
For a single chunk, S equals its logit. Identical scores produce the same value regardless of the number of chunks. A document label trains this shared aggregation; not every chunk in a risky script is automatically labeled positive. Gradients flow back to the chunks through the Smooth Max weights.
The advantage is complete input coverage with bounded memory requirements. The limitation is just as important: a chunk has no knowledge of other chunks. A variable value at the beginning of a script and its dangerous use many chunks later are not connected by document-wide attention.
5. Experiment II: Kestrel Expanded
The second hypothesis was simpler: perhaps the original linear approach mainly lacks suitable data coverage. We therefore retain character n-grams of length 3-5, a TF-IDF vectorizer with 50,000 features and a linear logistic classifier. We add examples from multiple shells, positive and negative PowerShell scripts, and synthetic obfuscation variants.
Kestrel Expanded is our experimental name for this extension of the model family. The concrete model is called linear_updated in the code; it is a newly trained Patronus model, not a new official Kestrel version. It processes the complete text rather than only a 510-byte prefix.
The additional fit set comprises 18,918 examples, including 10,672 positive and 8,246 negative cases. The existing original splits are preserved. Exact duplicates, contradictory identical commands and detected overlaps with held-out data are excluded. The expanded training set is also used in the encoder experiments.
What the data labels mean. The sources label examples such as
sudo systemctl stop apparmoras disabling a protection, andInvoke-Expression $env:PAYLOADas dynamic execution. They also labelcat /proc/net/tcpandnslookup evil.ioas data exfiltration. External data exfiltration cannot be inferred from these two commands alone. The labels are therefore a measurable target, not an independently confirmed malware verdict.
6. Results: context, obfuscation and thresholds
More context helps on long scripts
The PowerShell script test contains 1,000 positive and 1,000 negative files. All 2,000 scripts are longer than the former 510-byte window. The chunk encoder evaluates 349.22 MB of original text in 456,219 chunks; bytes omitted: zero. The test comprises the previously selected files up to 1 MiB.
Figure 2. F1 at thresholds selected separately on the original Bash validation set to optimize F1. The encoder's script score rises from 80.38% to 92.99%. Linear reaches 96.63%. The models are not operating at identical false-positive rates here.
The longer encoder variant achieves 89.50% recall with 97.00% specificity on scripts. Linear reaches 96.10% recall and 97.20% specificity. With the same new encoder and the same threshold, using only the first 1,022-byte chunk produces 88.71% F1; full aggregation achieves 92.99%. This internal comparison shows a benefit from the additional chunks.
Compared with the previous 510-byte model, F1 rises by 12.61 percentage points. This is not isolated evidence for chunking alone: window length, additional MLM adaptation, document training and pooling changed together. The paired bootstrap procedure groups scripts with identical earlier prefixes, but does not remove source or malware-family dependencies.
Obfuscation changes the representation
The paired tests transform the same original command: a Bash wrapper carries the text as Base64, quote variants change the spelling of the executable name, PowerShell variants use UTF-16/Base64, and CMD variants use caret escaping. This tests robustness to specific textual transformations, not every form of real-world obfuscation. Variants of the same original are not independent new attacks.
A strict threshold changes the picture
For a blocking mechanism, performance at low FPR is particularly relevant. We additionally investigate a shared threshold that allows at most 0.5% false positives in each of eight calibration groups. Earlier test cases are deterministically split into approximately 70% calibration and 30% holdout; variants follow their original. Linear and Encoder use exactly the same split and rule.
This analysis is exploratory: it consumes part of earlier test data, and possible thresholds were considered during the analysis. Its figures are not a new independent confirmation. The resulting cut is 1.805 for Linear and 7.46875 for the encoder.
Figure 3. Complete groups, including calibration cases. Darker recall cells indicate better detection; darker FPR cells indicate more false positives. Small PowerShell and CMD samples do not support a claim that the true FPR is zero.
The encoder performs slightly better than Linear on ordinary PowerShell and CMD commands. Linear is better on Bash Base64; recall is equal on Bash quotes. The most striking finding concerns PowerShell Base64 and long scripts: the encoder drops to 1.72% and 2.10% recall respectively, while Linear reaches 82.76% and 88.30%.
The encoder's binding calibration group is Bash Base64: its safe cases force the highest cut. The score scales of the other groups do not align well with it. At the lower original F1 cut, the encoder detects 96.55% of PowerShell Base64 cases but produces 9.84% false positives in that group. The model therefore contains detection signals; they do not yet provide sufficiently robust separation for a shared strict threshold.
Back to the original Bash test
| Model / threshold rule | Detections | False blocks | Recall | FPR | F1 |
|---|---|---|---|---|---|
| Expanded linear / strict exploratory rule | 142/193 | 6/4,001 | 73.6% | 0.15% | 83.3% |
| Chunk encoder / same rule | 130/193 | 1/4,001 | 67.4% | 0.025% | 80.2% |
| Kestrel / published per-case judgments | 178/193 | 10/4,001 | 92.2% | 0.25% | 93.4% |
Kestrel remains strong on its original task. At the cut shown, the encoder produces fewer false positives but detects considerably fewer positive cases. These rows compare different achieved FPRs and must not be read as a comparison at exactly the same operating point.
Expansion does not automatically preserve original performance either: our original linear Bash classifier achieved 92.39% F1 at its validation F1 threshold. The expanded model reaches 87.77% at its corresponding threshold, and the new chunk encoder reaches 86.09%. These internal Bash baselines are distinct from Kestrel. In this run, the transfer gain comes with a decline on the original test.
The comment padding test
We place the same held-out POSIX or PowerShell command after 4,096 or 16,384 bytes of repeated benign comment lines. Positive and negative cases receive identical padding. This additional test is neither trained on nor used to select the checkpoint, temperature or threshold; the models retain their original validation F1 thresholds.
Figure 4. Static robustness diagnostic with the unchanged command at the end of the document. Source labels are inherited; the scripts were not executed. The diagnostic applies to this specific form of repeated comment prefix.
Linear drops from 91.10% POSIX recall to 50.00% after 4 KiB and 41.78% after 16 KiB. On PowerShell, 56.90% and 51.72% remain respectively. The encoder falls to 0.68% POSIX recall after 4 KiB and to zero in the other padded cases. Both models achieve 100% observed specificity on these padded cases: they raise few alarms but miss many positive cases.
The encoder did not truncate the decisive bytes. Its failure concerns processing and aggregation. With normalized Smooth Max, many low-scoring chunks can reduce a single high chunk score; for the maximum z_max, z_max − τ log N ≤ S ≤ z_max. Changed context in the final chunk and a comment distribution different from training may also contribute. The contribution of each factor to this result has not yet been isolated.
For Linear, a related effect is plausible: repeated character patterns alter term frequencies and the document vector's L2 normalization. Relevant command features receive relatively less weight. This is an explanatory hypothesis, not a cause established by a separate ablation.
7. Conclusion: the data wins, the hard cases remain
The experiments support a pragmatic starting point: an expanded linear classifier is a very strong approach here. With suitable positive and negative data across multiple shell types, it handles transfer substantially better than the original Kestrel model in our extended evaluation. The small encoder does not outperform it overall.
The encoder experiment was still informative. Full chunking improves long scripts considerably. At the same time, the strict shared threshold shows how high F1 scores can hide differences between groups. The padding test reveals that complete input coverage does not guarantee reliable detection of isolated risks late in a document.
These results do not imply that linear models inherently understand Bash better or that encoders are unsuitable for shell risk. We compare specific models, sources and one training seed. Training volume, optimization, label quality and architecture are not fully separated. Small command groups and synthetic obfuscations further limit the conclusions.
The next useful experiments should isolate specific causes: compare unchanged chunk logits under different aggregations, vary comments and late payloads using training data only, and build a new untouched test from actual tool calls. Source and family splits, along with verified hard negatives, would be especially valuable. Shell-specific calibration may help, but would need independent validation before use.
Our result is not a victory for one architecture. It indicates which data and operating points a shell risk system actually needs to handle.
Sources and reproduction
- Kontext Security: ShellRisk-Bench. Task definition, original split and published per-case Kestrel judgments.
- Kontext Security: Kestrel. Official JSON model and published checksum. The version downloaded again on October 7 is documented in the experiment archive.
- Shell Safety. Additional POSIX, PowerShell and CMD command data. The source revision and fit/validation/test split are documented in the experiment protocol.
- Internal experiments: the cross-shell comparison of October 4, 2026, the chunk encoder run of October 5, and the threshold comparison and paired padding test of October 6. Raw scores, source revisions and data audits remain in the internal experiment directories and are not linked here as public downloads.
- The four figures were generated from stored results. The plotting script and underlying data belong to the experiment archive. No figure shows a threshold newly optimized on test data.
Study limitations. Public labels have not been independently audited. Exact deduplication and prefix checks do not establish independence of sources or malware families. The public PowerShell script collection supplies binary source labels, not malware verdicts confirmed in this experiment. We report static text classification. No dataset command or payload was executed.
All figures refer to the data and thresholds described here. The charts come from stored results, without synthetic performance values.