Logo loFT LLC

    • Website renewal and technical articles
    • Absorption merger of Lorchestra Inc.
    • IT Introduction Support Provider selection for FY2022
    • IT Introduction Support Provider selection for FY2021
    • Lorchestra Inc. established as a subsidiary
    • IT Introduction Support Provider selection for FY2020
    • loFT LLC established
      • Separating post-processing with Dagster and NATS JetStream
      • Moving PostgreSQL to always-on storage
      • Adding MLflow and MinIO to Dagster
      • Vector migration and configuration for a three-host homelab
      • PostgreSQL 18 and pgvector on rootless Quadlet
      • Planning an EPYC 9175F workstation
      • Placing data on NVMe/SATA and moving management UIs
      • QuteBrowser settings for keyboard-based monitoring
      • CRS304 VLANs and Linux egress control
      • Saving RouterOS counters during Syslog downtime
      • Podman and Quadlet: UIDs and DNS
      • Monitoring with Prometheus, Loki, and Quadlet
      • smartctl-exporter and rootful / rootless operating rules
      • A local development platform with EPYC and Podman
      • Quadlet operations on minimal Ubuntu
      • ELT memory allocation on one EPYC server
      • Backup and restore design for rootless Podman
      • A Compose stack for vLLM, llama.cpp, and a proxy
      • Moving Hugging Face models between cold and hot storage
      • Moving an OpenAI-compatible proxy from Rust to Go
      • An AI pipeline with Go, NATS and Dagster
      • familiar: a local LLM development platform and its observation tools
      • Moving llm-jp translation to on-demand batches
      • Task splitting and output_file in familiar
      • A Gemma 4 stack on two Blackwell GPUs
      • RAG and event processing in agent-gateway
      • agent-gateway v3: domain separation and MLflow integration
      • Embedding and rerank APIs with Rust and ONNX
      • Designing a WordPress-like Django blog
      • Generating a Django booking site with Qwen3.5
      • Jev-Omni multimodal decisions and comparison with Clef
      • Running DeepSeek Harness locally
      • DeepSeek V4 Flash 0731: vLLM measurements and CPU KV offload
      • Comparing GLM-5.2 GGUF quantization, MTP, and expert placement
      • DeepSeek-V4-Flash on two DwarfStar4 nodes
      • Local parallel-agent development with Step-3.7-Flash
      • Comparing vLLM and SGLang for Gemma 4 31B
      • MiMo V2.5 Pro on one and two GPUs
      • DeepSeek V4 Flash Q2 on DwarfStar 4
      • Qwen3.6-27B NVFP4, MTP, and LoRA measurements
      • Local DeepSeek-V4-Flash inference with llama.cpp and dual Blackwell GPUs
      • Qwen3.6-27B inference and role-specific LoRA plans
      • Kimi-K2.6 CPU / GPU placement and code generation
      • NVFP4 translation and three-style Japanese data
      • GLM-5.1 expert placement and hybrid inference
      • Django implementation and inference speed with Qwen3.5-397B-A17B
      • MiniMax-M2.7 speed and noncommercial limits
      • GLM-5.1 expert placement and generation speed
      • Managing conversation branches and evaluation data in Dagster
      • Generating a six-page dental clinic site with Qwen3.5
      • A CPU/GPU role plan for local LLMs
      • Bilingual system prompts for PLAMO
      • Designing 36 LTX-2 scenes and clip continuity
      • Hermes-4.3-36B in BF16, FP8 and nvfp4
      • IQuest-Coder-40B on CPU, GPU and Aider
      • Command A Reasoning in an Aider test-generation loop
      • A coding-assistance proposal with Serena MCP and Obsidian
      • Review checks with GLM-4.7-Flash Uncensored
      • IQuest-Coder Loop-Instruct generation speed in aider
      • MCP execution hosts and VSCode Remote SSH
      • Kimi-K2.5 on EPYC 9175F and a revised L3 cache hypothesis
      • MiniMax-2.5 Expert Offload and website generation
      • Qwen3.5-397B hybrid inference and CMS generation
      • Llama-4-Scout CPU / GPU throughput and caching
      • Kimi-K2.5 CPU inference and prompt cache
      • Llama 4 Maverick Q4 and Q8 CPU comparison
      • Three configurations for Qwen3-Coder-Next 80B
      • GLM-4.7-Flash on CPU, hybrid and GPU
      • DeepSeek-V3.2 throughput and cache mismatches
      • Static-site and Django CMS generation with Qwen3.5-397B
      • shelpa path restrictions, audit mirrors, and retirement
      • shelpa-mcp virtual pipelines and CWD management
      • Integrating voracle research into development
      • Nine Rust MCP servers in the homelab
      • From shelpa to filesystem: file operations and recovery
      • Obsidian search and LLM conversation import with voracle
      • Comparing accuracy after changing pathfinder MCP responses
      • Why aichat function calling hung with a symlinked tools directory
      • Path resolution and MCP checks in pathfinder
      • AST analysis and MCP integration in ctree
  • Articles
  • Profile
  • Photos
    Logo
    Contact Us
      • Japanese
    • to navigate
    • to select
    • to close
      • Home
      • Tech Memo
      • LLM Research
      On this page

      LLM Research

      Large language model benchmarks, CPU/GPU inference validation, and optimization research.

      These articles use AI-generated summaries of Obsidian notes originally kept as technical memos.
      English translations are produced with AI assistance.

      Generating a Django booking site with Qwen3.5

      A local web application development test with Qwen3.5-122B-A10B and MCP. It covers booking, admin, …

      Jev-Omni multimodal decisions and comparison with Clef

      AI integration tests of Jev-Omni decisions on 35 audio, image, Japanese email, and video inputs, …

      Running DeepSeek Harness locally

      DeepSeek-V4-Flash-Vision-Exp and V4.1-Flash ran on two RTX PRO 6000 Max-Q GPUs. This LLM integration …

      DeepSeek V4 Flash 0731: vLLM measurements and CPU KV offload

      DeepSeek V4 Flash 0731 on two RTX PRO 6000 96GB GPUs: DSpark K5 throughput, CPU KV offload, and a …

      Comparing GLM-5.2 GGUF quantization, MTP, and expert placement

      GLM-5.2 GGUF 1.630bpw and 2.244bpw on 192GB VRAM: speed, quality, low-VRAM placement, and measured …

      DeepSeek-V4-Flash on two DwarfStar4 nodes

      A local AI development test of IQ2XXS DeepSeek-V4-Flash on two GPUs. It covers throughput, DAG and …

      Local parallel-agent development with Step-3.7-Flash

      Six local agent roles with Step-3.7-Flash-NVFP4 generated a reservation API and admin system, with …

      Comparing vLLM and SGLang for Gemma 4 31B

      Gemma 4 31B NVFP4, FP8, and MTP on Blackwell 96GB: throughput, KV capacity, and coding-agent output …

      MiMo V2.5 Pro on one and two GPUs

      MiMo V2.5 Pro IQ2_S measured decode of 12.4–12.5 and 16.5–16.6 tok/s on one and two Blackwell 96GB …

      DeepSeek V4 Flash Q2 on DwarfStar 4

      DeepSeek V4 Flash 284B Q2-imatrix ran on one Blackwell 96GB GPU at 43.6 tok/s on short text and 31.4 …

      Qwen3.6-27B NVFP4, MTP, and LoRA measurements

      Qwen3.6-27B NVFP4 + MTP in vLLM: throughput, dynamic LoRA, VRAM budgeting, and tool errors for local …

      Local DeepSeek-V4-Flash inference with llama.cpp and dual Blackwell GPUs

      Local LLM deployment tests with DeepSeek-V4-Flash (284B MoE / 13B active), dual Blackwell Max-Q 96GB …

      Qwen3.6-27B inference and role-specific LoRA plans

      A one-day Qwen3.6-27B-FP8 run on SGLang measured about 100 tok/s singly and 160–180 tok/s combined …

      Kimi-K2.6 CPU / GPU placement and code generation

      Kimi-K2.6 IQ3_K and Q4_X in ik_llama.cpp: expert placement, local LLM throughput, and generated …

      NVFP4 translation and three-style Japanese data

      A 256-item translation and rewriting test with CAT-Translate and LLM-JP for AI training-data …

      GLM-5.1 expert placement and hybrid inference

      GLM-5.1 IQ3_KS measured 17–19 tok/s on two Blackwell 96GB GPUs and 768GB RAM. The local LLM …

      Django implementation and inference speed with Qwen3.5-397B-A17B

      A local LLM test for business application development. Qwen3.5-397B-A17B Q4_K_M (227.5 GiB) …

      MiniMax-M2.7 speed and noncommercial limits

      A noncommercial MiniMax-M2.7 evaluation on two Blackwell 96GB GPUs read an average 71.9 t/s from 10 …

      GLM-5.1 expert placement and generation speed

      A CPU/GPU placement test that raised GLM-5.1 IQ3_KS to 21.75 t/s. For AI system development …

      Managing conversation branches and evaluation data in Dagster

      Conversation replay, evaluation, and training datasets with Go, NATS, and Dagster: an AI development …

      Generating a six-page dental clinic site with Qwen3.5

      A local LLM test for website development: Qwen3.5 received a six-page dental site task using HTML, …

      A CPU/GPU role plan for local LLMs

      A local LLM integration plan separates async work and chat on Blackwell 96GB and EPYC 9175F. It …

      Bilingual system prompts for PLAMO

      Prompts for PLAMO-translate AI MODEL cover English-to-Japanese translation and Japanese pre-editing …

      Designing 36 LTX-2 scenes and clip continuity

      A video-generation AI workflow design separating scenarios, shots, and continuity. It covers fixed …

      Hermes-4.3-36B in BF16, FP8 and nvfp4

      Hermes-4.3-36B was compared in three formats on Blackwell 96GB and vLLM 0.14.0rc1. The LLM …

      IQuest-Coder-40B on CPU, GPU and Aider

      IQuest-Coder-V1-40B was measured with CPU Q5_K_M, GPU nvfp4 and Aider whole-edit. The LLM …

      Command A Reasoning in an Aider test-generation loop

      Go test generation with Command A Reasoning exposed Aider edit-format failures and a token-limit …

      A coding-assistance proposal with Serena MCP and Obsidian

      A proposal connecting Serena MCP, Obsidian, local LLMs, and VS Code, including the minimal setup and …

      Review checks with GLM-4.7-Flash Uncensored

      GLM-4.7-Flash Uncensored was evaluated for Rust-code review and throughput. It can suggest review …

      IQuest-Coder Loop-Instruct generation speed in aider

      IQuest-Coder-V1-40B-Loop-Instruct generated at 0.6–8 tok/s in aider. For local AI coding setup, we …

      MCP execution hosts and VSCode Remote SSH

      In the tested setup, Zed launched MCP on Mac while VSCode Remote SSH used the compute runtime. This …

      Kimi-K2.5 on EPYC 9175F and a revised L3 cache hypothesis

      Kimi-K2.5 on EPYC 9175F with 768GB RAM: thread counts, 128K context, prompt caching, and a revised …

      MiniMax-2.5 Expert Offload and website generation

      MiniMax-2.5 IQ5_K through IQ3_S: Expert Offload for local LLM integration and React landing-page and …

      Qwen3.5-397B hybrid inference and CMS generation

      28 Qwen3.5-397B IQ4_NL runs on EPYC and GPUs, followed by Django CMS scaffolding. Throughput, …

      Llama-4-Scout CPU / GPU throughput and caching

      Llama-4-Scout CPU Q6_K and GPU nvfp4 measurements: throughput, caching, and context limits for batch …

      Kimi-K2.5 CPU inference and prompt cache

      Kimi-K2.5 Q4_K_S/Q4_K_M was measured on EPYC 9175F. The async LLM integration check covers th=13 …

      Llama 4 Maverick Q4 and Q8 CPU comparison

      A Q4_K_M versus Q8_0 test on EPYC 9175F with 768GB memory. For LLM integration choices, it records …

      Three configurations for Qwen3-Coder-Next 80B

      Qwen3-Coder-Next was measured in BF16 CPU, IQ4_NL hybrid/GPU, and nvfp4 modes for local coding …

      GLM-4.7-Flash on CPU, hybrid and GPU

      GLM-4.7-Flash IQ5_K measured PP of 100/1635/3723 tok/s and TG of 20/67/99 tok/s. This local LLM …

      DeepSeek-V3.2 throughput and cache mismatches

      Local DeepSeek-V3.2 Speciale logs reviewed for LLM integration performance. The record covers PP/TG, …

      Static-site and Django CMS generation with Qwen3.5-397B

      A six-page static site and Django CMS generated from specifications with Qwen3.5-397B, including …

      © 2017-2026 loFT LLC

      Privacy Policy Security Whitepaper