Hi, I'm Shane 👋

Software Engineer interested in C++, AI Infra and LLM Systems.

记录技术与项目,也记录一些长期值得留下的东西。

GitHub · RSS

Recent Writing
View all →
Selected Projects
View all →
  • cuflash

    从零手写的 CUDA FlashAttention 实现:标量内核到 WMMA Tensor Core 前向、FlashDecoding/Split-KV 与 Roofline 性能分析,支持 FP16/BF16/FP32 前反向,覆盖 sm_70–sm_90。

    CUDA · C++ · FlashAttention · Tensor Core

  • tiny-llm

    CUDA/C++17 精简 LLM 推理运行时:GGUF 加载与反量化、W8A16 推理、显式与分页 KV Cache、tokenizer、采样,以及 CUDA Graphs 加速 decode。

    CUDA · C++ · LLM Inference · KV Cache

  • paged-serving

    Rust 实现的 LLM Serving 控制面:Paged KV 调度、continuous batching、OpenAI 兼容 API 与 SSE,经 C ABI 接入 tiny-llm 真实 CUDA 后端。

    Rust · LLM Serving · Systems

Experience
More →
  • 2022 — Present
    BGI 华大基因
    Software Engineer
    Bioinformatics HPC · AI Platform
  • 2021 — 2022
    即构科技 ZEGO
    Backend Software Engineer
    实时音视频后台开发
  • 2018 — 2020
    迈瑞医疗
    Software Engineer · Medical Algorithm Software Development
    医学算法软件开发
Connect

欢迎通过以下方式联系我,交流技术或随便聊聊。