Linpeng Tang
- 🚀 Chief AI Scientist @ SpacemiT | 🎓 Princeton CS PhD |
Moqi Co-founder & CTO | Meta Systems - 👑 Technical Leadership: Interdisciplinary R&D leadership and strategic vision, bridging algorithms, systems, and hardware co-design to deliver high-impact applications

Hi, I am an AI researcher, engineer, and technical leader, currently serving as the Chief AI Scientist at SpacemiT, leading the RISC-V AI software foundation and open ecosystem development. I have long been dedicated to the deep integration and co-design of AI algorithms and high-performance underlying systems: from large-scale distributed storage and caching systems at Meta (Facebook), to billion-scale biometric recognition and edge sensing, to the MyScaleDB AI database with SQL-vector synergy, and next-generation data infrastructure for LLMs and agents. Currently, my primary mission is to leverage the open RISC-V AI architecture to bridge custom instruction extensions, compiler optimization, operator acceleration, and edge LLM / Physical AI deployment, establishing an open, flexible, and high-efficiency software-hardware co-design foundation for AI computing.
I hold a Ph.D. in Computer Science from Princeton University, advised by Prof. Kai Li. My work has been recognized with honors including the WAIC (World AI Conference) SAIL Award and 1st place in the KDDCup.
Technical Thoughts
Selected Recognition
- 🏆 WAIC (World AI Conference) SAIL Award, 2024
- 🥇 First Prize, HICOOL Global Entrepreneur Summit, 2022
- ⚙️ Flagship Systems: RISC-V AI Software Ecosystem, Data-Centric AI Platform, MyScale AI Database,
Billion-scale national fingerprint database - 📚 Top-tier Publications: NSDI, KDD, FAST, CIKM Best Student Paper,
KDDCup 1st Place
Experience
SpacemiT
Chief AI Scientist | 2026 - Present
Institute of Advanced Algorithms Research
Data&AI Center | 2024 - 2026
Moqi Technology
Co-founder & CTO | 2017 – 2024
Meta
Research Consultant | 2013 – 2016
HP Labs Beijing
Research Intern | 2011 - 2012
Education
Princeton University
Ph.D. in Computer Science | 2012 – 2018
Advisor: Prof. Kai Li (Member of National Academy of Engineering)
Shanghai Jiao Tong University
B.S. in Computer Science, ACM Class | 2008 – 2012
Products & Projects
RISC-V AI Software Stack & Open Ecosystem 2026 – Present
- Full-Stack AI Software & Compiler Infrastructure: Building the AI software foundation targeting RISC-V Vector and Tensor AI computation units, spanning AI compilers, high-performance operator libraries, low-bit quantization, runtime engines, and hardware-software co-design iterations.
- Edge LLMs & Embodied AI Deployment: Driving hardware-software co-optimization for lightweight LLMs/VLMs, vision, and embodied / Physical AI models on RISC-V chip platforms, achieving high energy efficiency and low-latency on-device inference.
- Open Ecosystem & Standardization: Deeply connecting mainstream open-source ecosystems (PyTorch, vLLM, LLVM, etc.) with the RISC-V community, promoting RISC-V as an open standard and mainstream choice for AI computing.
Data-Centric AI Platform 2024 – 2026
- Led overall product architecture design and key project delivery, guiding the team to build a new generation of Agentic LLM & Data infrastructure.
- Pioneered the development and implementation of a multimodal data intelligence pipeline system based on agents and the DataFlow data preparation framework. Built-in with 150+ intelligent operators, it supports natural language conversational automated pipeline orchestration, enabling highly efficient and flexible processing of massive heterogeneous data.
- Addressing the high-risk hallucination challenge of LLMs in scientific and industrial scenarios, constructed a high-fidelity data synthesis and feedback system based on multi-tier verifiers (including rule filtering, knowledge graphs and simulations).
- Disrupted the traditional data engineering paradigm that consumes 90% human effort, significantly lowering the barrier to producing AI-ready datasets. Successfully deployed in multiple benchmark scenarios such as industrial manufacturing, multimodal corpus management, and scientific corpora, drastically reducing the costs for enterprises to build specialized agents and LLMs.
MyScaleDB AI Database 2020 – Present
- Responsible for defining product technical architecture, leading core vector search algorithm design and core engine R&D, creating a world-leading open-source AI database system.
- Pioneered the concept of an AI database in the industry, innovatively achieving integrated management and joint retrieval of PB-level structured and unstructured data (vectors, graphs, text, spatio-temporal, etc.) within a single SQL kernel based on a columnar data engine.
- Self-developed the MSTG vector engine and deeply combined it with a high-performance NVMe SSD memory caching mechanism for software-hardware co-optimization. While ensuring millisecond-level complex joint queries, achieved a 10x increase in vector data storage density.
- Successfully implemented in large-scale knowledge base constructions for industrial manufacturing, AI for Science, and financial auxiliary decision-making, providing exceptional cost-effectiveness for massive corpora and widely used in a global SaaS.
Contactless Fingerprint & Palmprint Capture Device 2018 – 2022
- Led the product definition of the world's first large-area, high-quality contactless fingerprint and palmprint capture terminal. Guided the team to overcome core technical challenges such as 3D reconstruction and complex optical image enhancement.
- Combining binocular vision with a self-developed structured light system, achieved sub-millimeter high-precision 3D reconstruction of fingers. Introduced multi-source, multi-band optical designs and deep learning image enhancement algorithms, substantially breaking through ambient light interference.
- Successfully disrupted industry pain points and technical bottlenecks of traditional contact-based capture, launching revolutionary contactless capture terminal devices, and driving inter-generational technological upgrades in security biometric capture hardware.
Massive Fingerprint Identification System 2015 – 2022
- Responsible for core system architecture design and deep learning model R&D for massive fingerprint and palmprint matching.
- Pioneered a multi-scale vector representation scheme, innovatively introducing an Active Deep Learning mechanism to drive model self-optimization and iteration. With joint CPU and GPU acceleration, broke through the technical bottleneck of 100-billion scale multi-scale feature indexing.
- Improved the speed, accuracy, and automation of massive complex biometric feature retrieval by over 100 times. Successfully deployed at the National Fingerprint Center, generating significant social impact.
Video Popularity Prediction System 2015 – 2017
- Responsible for high-performance algorithm design and implementation for Facebook's massive-scale video traffic trend prediction.
- Self-developed a high-performance time-series probabilistic prediction model, deeply coupling and optimizing it with the underlying video compression strategy flow and real-time cache scheduling pipeline.
- Achieved real-time accurate prediction of large-scale video popularity, improving prediction accuracy by over 10%. Supported Facebook in adopting smarter video compression conversion schemes and efficient cache scheduling, reducing system consumption while enhancing user viewing experience on the platform.
RIPQ Caching System 2013 – 2015
- Responsible for core algorithm design and system implementation of a large-scale cache scheduling system based on SSD storage.
- Pioneered the Restricted Insertion Priority Queue (RIPQ) caching algorithm, cleverly resolving the inherent non-sequential write amplification and sharp performance drop issues of traditional cache eviction mechanisms on Solid State Drives (SSDs) from the bottom layer.
- Built a next-generation intelligent caching system with extremely low write amplification and high throughput features. Successfully deployed in Facebook's global CDN edge nodes and core caching systems, increasing cache hit rates by over 20% in large-scale concurrent environments, optimizing network request latency, and saving massive bandwidth costs.