Review checks with GLM-4.7-Flash Uncensored
GLM-4.7-Flash Uncensored was evaluated for Rust-code review and throughput. It can suggest review angles for LLM-assisted development, but people must verify vulnerability preconditions.
Checked outputs
I asked uncensored GLM-4.7 Flash for attack-analysis and reproduction ideas on Rust code. I checked the logs and throughput for its suitability in defensive review.
It identified functions and test candidates but also asserted vulnerabilities without checking preconditions. I would limit it to idea generation in this workload.
Review candidate generation
The run showed these useful behaviors:
Useful observations
- It generated candidate attack scenarios and test inputs.
- It identified functions to inspect, including
to_rel_string,collect_files,build_globset, andsanitize_symbol_text. - I consider it useful for initial threat modeling, review questions and candidate test inputs.
These outputs can broaden the candidates considered at the start of a review.
Missed preconditions
The following claims lacked precondition checks:
- It called
strip_prefix + unwrap_orpath traversal even though, in isolation, it only falls back to returning an absolute path for display. - It explained Rust
regexin terms of catastrophic backtracking, even though the crate is generally built around linear-time execution. - It described
globproblems as “injection” when the real issues are usually boundary control and IO or computational blowups. - It linked
sanitize_symbol_textdirectly to XSS without proving the rendering-side precondition.
A relevant function does not establish a valid finding. Data flow and boundaries must be verified before accepting the conclusion.
Suitability for defensive security work
For defensive work, I would not use this model as evidence.
- The model may generate exploit-oriented code.
- Logs and prompt histories retain that output.
- Findings and generated artifacts both require review and handling rules.
The proposed use is isolated local brainstorming, with generated material kept private.
Scope of review assistance
The role division is:
- Let it do idea generation, candidate test generation, and review coverage.
- Do not let it make vulnerability claims, CVE-grade assertions, or direct exploit plans that are accepted without review.
- Require human verification of data flow, trust boundaries, permission models, and caller-side input constraints.
Human verification remains part of development-assistance LLM integration.
Performance observations
I also recorded inference settings and speed.
Model and runtime conditions
- GGUF:
Q8_0 - model params:
29.943B - model size:
29.924GiB (8.584 BPW) n_ctx = 131072n_batch = 2048n_ubatch = 2048flash_attn = 1fused_moe = 1mla_attn = 3- GPU:
NVIDIA RTX PRO 6000 Blackwell Max-Q 96GB - layer offload:
48/48 layers GPU - KV cache:
CUDA0 KV buffer size = 3595.52MiB - compute buffer:
CUDA0 compute buffer 7360.62MiB - host compute buffer:
CUDA_Host compute buffer 528.02MiB - CPU buffer:
28152.00MiB
Some log entries suggest CPU-resident expert weights. 48/48 layer offload alone does not prove that every weight is on GPU.
Representative throughput
Representative log values:
prompt eval: 6194.54 ms / 9519 tokens = 1536.68 tok/seval: 1657.83 ms / 111 tokens = 66.95 tok/sprompt eval: 949.06 ms / 125 tokens = 131.71 tok/seval: 6837.71 ms / 232 tokens = 33.93 tok/s
Observed decode was roughly 34-67 tok/s, varying by context, cache hits and request.
Benchmark details
Per-request measurements are shown below.
| # | PP(tok) | TG(tok) | Ctx_used | T_PP(s) | S_PP(t/s) | T_TG(s) | S_TG(t/s) | total(s) |
|---|---|---|---|---|---|---|---|---|
| 1 | 125 | 232 | 357 | 0.949 | 131.71 | 6.838 | 33.93 | 7.787 |
| 2 | 661 | 430 | 1091 | 3.295 | 200.63 | 12.852 | 33.46 | 16.147 |
| 3 | 467 | 404 | 871 | 2.555 | 182.77 | 12.006 | 33.65 | 14.561 |
| 4 | 783 | 448 | 1231 | 3.755 | 208.50 | 13.573 | 33.01 | 17.329 |
| 5 | 761 | 400 | 1161 | 3.705 | 205.38 | 12.102 | 33.05 | 15.807 |
| 6 | 916 | 410 | 1326 | 4.179 | 219.19 | 12.499 | 32.80 | 16.678 |
| 7 | 839 | 512 | 1351 | 3.996 | 209.95 | 15.790 | 32.43 | 19.786 |
| 8 | 497 | 512 | 1009 | 2.652 | 187.43 | 15.755 | 32.50 | 18.407 |
| 9 | 9519 | 111 | 9630 | 6.195 | 1536.68 | 1.658 | 66.95 | 7.852 |
| 10 | 11676 | 525 | 12201 | 8.860 | 1317.88 | 8.403 | 62.48 | 17.263 |
| 11 | 12291 | 66 | 12357 | 7.754 | 1585.20 | 0.959 | 68.84 | 8.712 |
| 12 | 532 | 551 | 1083 | 1.365 | 389.62 | 8.819 | 62.48 | 10.184 |
| 13 | 558 | 777 | 1335 | 1.408 | 396.17 | 12.473 | 62.30 | 13.881 |
| 14 | 784 | 778 | 1562 | 1.471 | 532.92 | 12.478 | 62.35 | 13.949 |
| 15 | 784 | 732 | 1516 | 1.484 | 528.15 | 11.804 | 62.01 | 13.288 |
| 16 | 739 | 1781 | 2520 | 1.502 | 491.91 | 29.067 | 61.27 | 30.570 |
| 17 | 1792 | 626 | 2418 | 1.764 | 1015.92 | 10.070 | 62.17 | 11.834 |
Ctx_used is approximated as PP+TG, not cumulative n_past. The logs did not have a separate total-token field.
Throughput and decision quality
A separate aggregate was recorded as GPU full offload:
- Total tokens:
64,404 - Prompt:
33,295 - Generation:
31,109 - Total time:
160.487s - Prompt time:
9.089s - Generation time:
151.304s - Weighted average throughput:
Prompt 3,663.2 tok/s - Weighted average throughput:
Generation 205.6 tok/s - Fastest generation:
req_id=2316 (511.3 tok/s, 247 tok / 0.484s) - Slowest generation:
req_id=4522 (103.2 tok/s, 1963 tok / 19.028s) - Fastest prompt:
req_id=0 (7009.3 tok/s, 5309 tok / 0.757s) - Slowest prompt:
req_id=2261 (662.5 tok/s, 104 tok / 0.157s)
| req_id | prompt_tokens | gen_tokens | total_tokens | ctx_tokens | T_PP(s) | T_TG(s) | total(s) | TPS_PP | TPS_TG | ms/token_PP | ms/token_TG |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 5309 | 630 | 5939 | 5939 | 0.757 | 4.852 | 5.610 | 7009.3 | 129.8 | 0.14 | 7.70 |
| 632 | 1089 | 742 | 1831 | 1831 | 0.192 | 5.798 | 5.990 | 5659.7 | 128.0 | 0.18 | 7.81 |
| 1375 | 2369 | 325 | 2694 | 2694 | 0.416 | 2.305 | 2.721 | 5701.1 | 141.0 | 0.18 | 7.09 |
| 1701 | 756 | 400 | 1156 | 1156 | 0.175 | 3.166 | 3.342 | 4307.9 | 126.3 | 0.23 | 7.92 |
| 2102 | 444 | 70 | 514 | 514 | 0.149 | 0.550 | 0.699 | 2971.8 | 127.4 | 0.34 | 7.85 |
| 2173 | 246 | 87 | 333 | 333 | 0.139 | 0.686 | 0.825 | 1767.5 | 126.8 | 0.57 | 7.89 |
| 2261 | 104 | 54 | 158 | 158 | 0.157 | 0.424 | 0.581 | 662.5 | 127.5 | 1.51 | 7.84 |
| 2316 | 65 | 247 | 312 | 312 | 0.127 | 1.968 | 2.095 | 511.3 | 125.5 | 1.96 | 7.97 |
| 2564 | 913 | 118 | 1031 | 1031 | 0.212 | 0.942 | 1.155 | 4296.8 | 125.2 | 0.23 | 7.99 |
| 2683 | 135 | 227 | 362 | 362 | 0.133 | 1.815 | 1.948 | 1011.6 | 125.1 | 0.99 | 8.00 |
| 2911 | 1546 | 244 | 1790 | 1790 | 0.318 | 1.989 | 2.308 | 4855.5 | 122.7 | 0.21 | 8.15 |
| 3156 | 1824 | 320 | 2144 | 2144 | 0.375 | 2.656 | 3.031 | 4859.6 | 120.5 | 0.21 | 8.30 |
| 3477 | 5230 | 365 | 5595 | 5595 | 1.414 | 3.435 | 4.849 | 3698.3 | 106.3 | 0.27 | 9.41 |
| 3844 | 1531 | 401 | 1932 | 1932 | 0.553 | 3.788 | 4.340 | 2770.8 | 105.9 | 0.36 | 9.45 |
| 4246 | 779 | 275 | 1054 | 1054 | 0.403 | 2.597 | 3.001 | 1932.0 | 105.9 | 0.52 | 9.45 |
| 4522 | 1357 | 1963 | 3320 | 3320 | 0.540 | 19.028 | 19.568 | 2510.9 | 103.2 | 0.40 | 9.69 |
| 6486 | 1977 | 2048 | 4025 | 4025 | 0.631 | 20.136 | 20.767 | 3135.1 | 101.7 | 0.32 | 9.83 |
| 8535 | 2063 | 2048 | 4111 | 4111 | 0.915 | 20.204 | 21.119 | 2255.2 | 101.4 | 0.44 | 9.87 |
| 10584 | 2058 | 1879 | 3937 | 3937 | 0.928 | 18.442 | 19.370 | 2217.3 | 101.9 | 0.45 | 9.81 |
| 12464 | 1891 | 1689 | 3580 | 3580 | 0.734 | 16.605 | 17.339 | 2575.2 | 101.7 | 0.39 | 9.83 |
This aggregate is kept separate from the preceding placement measurements.
Risk of accepting incorrect findings
Throughput does not establish security-decision accuracy.
Claims without proven preconditions still need individual validation, even when long inputs run quickly.
A shared environment also needs rules for logs and histories containing attack-related code.
Instructions missing precondition checks
I plan to change the Rust-review instructions as follows.
Instructions to test next
give me exploit code
That instruction alone omits the precondition-checking steps.
Scope of use
- Ask for
input source -> boundary -> sinkdata flow first. - Ask for attack preconditions and require the model to say which ones are not yet proven.
- Ask for safe minimal reproduction tests using non-dangerous synthetic input.
- Ask for false-positive explanations.
- Ask for remediation guidance that includes trust boundaries and caller-side constraints.
Whether these instructions reduce false claims remains to be tested.
Results and next steps
The candidate uses are:
- Attack-analysis ideation.
- Review assistance with human precondition verification.
- Throughput checks for long inputs and cache reuse.
