No description
  • Python 90.9%
  • Cuda 4.3%
  • C++ 2.3%
  • C 1.9%
  • Shell 0.6%
Find a file
centra a72b8e3979 feat(qwen38): preserve centra inference and stability improvements
Snapshot the local Qwen3.8-Flash-Next implementation based on upstream 58f4b9e: file-backed PLE, QSA KV quantization, NVFP4 and resident FP8 loading, MTP graphs and sampling, layer LRU, multimodal serving, and cancellation fixes. Include existing experimental paths and regression tests, plus fork documentation.

Validation: 166 targeted CPU tests passed, 1 skipped; syntax checks for 122 changed Python files; git diff --check. GPU inference and full performance tests were not rerun for this snapshot.
2026-09-28 14:48:46 +09:00
.github/workflows build(release): pypi-ready wheels, metadata, and publish workflow 2026-08-18 05:00:47 +00:00
assets chore(assets): update wechat group QR code 2026-08-26 00:41:24 -07:00
benchmarks feat(qwen38): preserve centra inference and stability improvements 2026-09-28 14:48:46 +09:00
docs feat(qwen38): preserve centra inference and stability improvements 2026-09-28 14:48:46 +09:00
freetoken-kernel-cache build(kernel-cache): add sm_80 (A100/A800) to default arches (#75) 2026-08-22 23:38:42 -07:00
python/freetoken feat(qwen38): preserve centra inference and stability improvements 2026-09-28 14:48:46 +09:00
scripts feat(qwen38): preserve centra inference and stability improvements 2026-09-28 14:48:46 +09:00
tests feat(qwen38): preserve centra inference and stability improvements 2026-09-28 14:48:46 +09:00
.gitignore feat: initial open-source release 2026-08-11 22:53:25 +00:00
CONTRIBUTING.md docs: add CONTRIBUTING.md 2026-08-22 22:22:15 -07:00
install.sh build(release): pypi-ready wheels, metadata, and publish workflow 2026-08-18 05:00:47 +00:00
LICENSE feat: initial open-source release 2026-08-11 22:53:25 +00:00
pyproject.toml build(release): pypi-ready wheels, metadata, and publish workflow 2026-08-18 05:00:47 +00:00
README.md feat(qwen38): preserve centra inference and stability improvements 2026-09-28 14:48:46 +09:00
setup.py feat: initial open-source release 2026-08-11 22:53:25 +00:00

Centra fork: Qwen3.8-Flash-Nextを単一RTX 5090で運用するためのローカル改修版です。 改修内容・使用構成・検証範囲は Centra fork notes を参照してください。 このforkを使う場合は、このリポジトリからソースインストールしてください。以下の上流向けPyPIインストールでは改修は入りません。

FreeToken

| Download | Paper | Developer Slack | Community Discord | Community WeChat |

Unlock datacenter-class intelligence on the hardware you already own — Run 290B+ frontier MoE models locally on your gaming PC at blistering interactive speeds.

About

FreeToken is an edge-native Mixture-of-Experts (MoE) serving engine designed for running frontier-scale open-weight models on personal and consumer hardware. It treats heterogeneous edge resources—GPUs, CPUs, host memory, and interconnects—as a unified, elastic inference platform. Its core features include:

  • Fast Edge-Native Runtime: Provides efficient MoE serving with bandwidth-adaptive CPU–GPU co-execution (q^\star policy), full-layer double-buffered prefill streaming, global LRU expert caching, graph-compatible execution, and the FTW fast weight format.
  • Semantic-Aware Caching: Features semantic anchor checkpoints for recurrent state and KV caches, allowing agentic context edits (e.g., tool calls, thinking blocks) to avoid redundant context recomputation.
  • Elastic Memory Management: Supports dynamic, runtime VRAM re-allocation between expert caches and KV memory without engine restarts or weight reloading.
  • Broad MoE & Ecosystem Support: Supports frontier open-weight MoE models (e.g., DeepSeek-V4-Flash, Qwen3.6-35B-A3B, GLM-5.2) across various parameter scales and quantization formats (e.g., MXFP4, NVFP4, FP8, BF16), with Anthropic/OpenAI-compatible APIs for seamless integration with real-world coding and tool-calling agents (e.g., Codex, Claude Code, OpenCode, OpenClaw, DeepSeek Harness).
  • Diverse Consumer Hardware: Scales across consumer laptops, gaming desktops, and workstation GPUs, with native support for NVIDIA RTX 30, RTX 40, and RTX 50 series GPUs.

Getting Started

Desktop app

Download FreeToken for Windows or Linux at flashml.ai. It sets the engine up for you and gives you a GUI for running models, chatting, and tuning the engine.

FreeToken Desktop

CLI

Install FreeToken with uv (recommended) or pip:

uv pip install "freetoken[accel]"

Or build from source:

git clone https://github.com/FlashML-org/FreeToken.git && cd FreeToken
uv venv && source .venv/bin/activate
uv pip install -e ".[accel]"

For More details:

Citation

If you use FreeToken for your research, please cite our paper:

@article{yang2026freetoken,
  title={FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution},
  author={Yang, Shuo and Fan, Xiaoze and Pan, Melissa and Xi, Haocheng and Wang, Zhe and Sun, Shanlin and Keutzer, Kurt and Han, Song and Zaharia, Matei and Xu, Chenfeng and Stoica, Ion},
  journal={arXiv preprint arXiv:2608.16157},
  year={2026}
}

Acknowledgment

FreeToken was deeply inspired by mini-sglang, and learned the design and reused code from the following projects: SGLang, vLLM, FlashInfer, flash-linear-attention, LightLLM and llama.cpp.

License

Apache License 2.0.