Checked outputs

I asked uncensored GLM-4.7 Flash for attack-analysis and reproduction ideas on Rust code. I checked the logs and throughput for its suitability in defensive review.

It identified functions and test candidates but also asserted vulnerabilities without checking preconditions. I would limit it to idea generation in this workload.

Review candidate generation

The run showed these useful behaviors:

Useful observations

  • It generated candidate attack scenarios and test inputs.
  • It identified functions to inspect, including to_rel_string, collect_files, build_globset, and sanitize_symbol_text.
  • I consider it useful for initial threat modeling, review questions and candidate test inputs.

These outputs can broaden the candidates considered at the start of a review.

Missed preconditions

The following claims lacked precondition checks:

  • It called strip_prefix + unwrap_or path traversal even though, in isolation, it only falls back to returning an absolute path for display.
  • It explained Rust regex in terms of catastrophic backtracking, even though the crate is generally built around linear-time execution.
  • It described glob problems as “injection” when the real issues are usually boundary control and IO or computational blowups.
  • It linked sanitize_symbol_text directly to XSS without proving the rendering-side precondition.

A relevant function does not establish a valid finding. Data flow and boundaries must be verified before accepting the conclusion.

Suitability for defensive security work

For defensive work, I would not use this model as evidence.

  • The model may generate exploit-oriented code.
  • Logs and prompt histories retain that output.
  • Findings and generated artifacts both require review and handling rules.

The proposed use is isolated local brainstorming, with generated material kept private.

Scope of review assistance

The role division is:

  • Let it do idea generation, candidate test generation, and review coverage.
  • Do not let it make vulnerability claims, CVE-grade assertions, or direct exploit plans that are accepted without review.
  • Require human verification of data flow, trust boundaries, permission models, and caller-side input constraints.

Human verification remains part of development-assistance LLM integration.

Performance observations

I also recorded inference settings and speed.

Model and runtime conditions

  • GGUF: Q8_0
  • model params: 29.943B
  • model size: 29.924GiB (8.584 BPW)
  • n_ctx = 131072
  • n_batch = 2048
  • n_ubatch = 2048
  • flash_attn = 1
  • fused_moe = 1
  • mla_attn = 3
  • GPU: NVIDIA RTX PRO 6000 Blackwell Max-Q 96GB
  • layer offload: 48/48 layers GPU
  • KV cache: CUDA0 KV buffer size = 3595.52MiB
  • compute buffer: CUDA0 compute buffer 7360.62MiB
  • host compute buffer: CUDA_Host compute buffer 528.02MiB
  • CPU buffer: 28152.00MiB

Some log entries suggest CPU-resident expert weights. 48/48 layer offload alone does not prove that every weight is on GPU.

Representative throughput

Representative log values:

  • prompt eval: 6194.54 ms / 9519 tokens = 1536.68 tok/s
  • eval: 1657.83 ms / 111 tokens = 66.95 tok/s
  • prompt eval: 949.06 ms / 125 tokens = 131.71 tok/s
  • eval: 6837.71 ms / 232 tokens = 33.93 tok/s

Observed decode was roughly 34-67 tok/s, varying by context, cache hits and request.

Benchmark details

Per-request measurements are shown below.

#PP(tok)TG(tok)Ctx_usedT_PP(s)S_PP(t/s)T_TG(s)S_TG(t/s)total(s)
11252323570.949131.716.83833.937.787
266143010913.295200.6312.85233.4616.147
34674048712.555182.7712.00633.6514.561
478344812313.755208.5013.57333.0117.329
576140011613.705205.3812.10233.0515.807
691641013264.179219.1912.49932.8016.678
783951213513.996209.9515.79032.4319.786
849751210092.652187.4315.75532.5018.407
9951911196306.1951536.681.65866.957.852
1011676525122018.8601317.888.40362.4817.263
111229166123577.7541585.200.95968.848.712
1253255110831.365389.628.81962.4810.184
1355877713351.408396.1712.47362.3013.881
1478477815621.471532.9212.47862.3513.949
1578473215161.484528.1511.80462.0113.288
16739178125201.502491.9129.06761.2730.570
17179262624181.7641015.9210.07062.1711.834

Ctx_used is approximated as PP+TG, not cumulative n_past. The logs did not have a separate total-token field.

Throughput and decision quality

A separate aggregate was recorded as GPU full offload:

  • Total tokens: 64,404
  • Prompt: 33,295
  • Generation: 31,109
  • Total time: 160.487s
  • Prompt time: 9.089s
  • Generation time: 151.304s
  • Weighted average throughput: Prompt 3,663.2 tok/s
  • Weighted average throughput: Generation 205.6 tok/s
  • Fastest generation: req_id=2316 (511.3 tok/s, 247 tok / 0.484s)
  • Slowest generation: req_id=4522 (103.2 tok/s, 1963 tok / 19.028s)
  • Fastest prompt: req_id=0 (7009.3 tok/s, 5309 tok / 0.757s)
  • Slowest prompt: req_id=2261 (662.5 tok/s, 104 tok / 0.157s)
req_idprompt_tokensgen_tokenstotal_tokensctx_tokensT_PP(s)T_TG(s)total(s)TPS_PPTPS_TGms/token_PPms/token_TG
05309630593959390.7574.8525.6107009.3129.80.147.70
6321089742183118310.1925.7985.9905659.7128.00.187.81
13752369325269426940.4162.3052.7215701.1141.00.187.09
1701756400115611560.1753.1663.3424307.9126.30.237.92
2102444705145140.1490.5500.6992971.8127.40.347.85
2173246873333330.1390.6860.8251767.5126.80.577.89
2261104541581580.1570.4240.581662.5127.51.517.84
2316652473123120.1271.9682.095511.3125.51.967.97
2564913118103110310.2120.9421.1554296.8125.20.237.99
26831352273623620.1331.8151.9481011.6125.10.998.00
29111546244179017900.3181.9892.3084855.5122.70.218.15
31561824320214421440.3752.6563.0314859.6120.50.218.30
34775230365559555951.4143.4354.8493698.3106.30.279.41
38441531401193219320.5533.7884.3402770.8105.90.369.45
4246779275105410540.4032.5973.0011932.0105.90.529.45
452213571963332033200.54019.02819.5682510.9103.20.409.69
648619772048402540250.63120.13620.7673135.1101.70.329.83
853520632048411141110.91520.20421.1192255.2101.40.449.87
1058420581879393739370.92818.44219.3702217.3101.90.459.81
1246418911689358035800.73416.60517.3392575.2101.70.399.83

This aggregate is kept separate from the preceding placement measurements.

Risk of accepting incorrect findings

Throughput does not establish security-decision accuracy.

Claims without proven preconditions still need individual validation, even when long inputs run quickly.

A shared environment also needs rules for logs and histories containing attack-related code.

Instructions missing precondition checks

I plan to change the Rust-review instructions as follows.

Instructions to test next

  • give me exploit code

That instruction alone omits the precondition-checking steps.

Scope of use

  • Ask for input source -> boundary -> sink data flow first.
  • Ask for attack preconditions and require the model to say which ones are not yet proven.
  • Ask for safe minimal reproduction tests using non-dangerous synthetic input.
  • Ask for false-positive explanations.
  • Ask for remediation guidance that includes trust boundaries and caller-side constraints.

Whether these instructions reduce false claims remains to be tested.

Results and next steps

The candidate uses are:

  • Attack-analysis ideation.
  • Review assistance with human precondition verification.
  • Throughput checks for long inputs and cache reuse.