Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Mini-LLM-Server

Minimal server for local LLMs equipped with several inference optimization techniques that can be configured via a UI.

Project is still under development. Few optimizations have been added, few more to come.

Features (some are TODO)

  • Batching (static and dynamic)
  • KV Caching
  • Prompt Caching (Prefix Caching)
  • Tensor Parallelism
  • Speculative Decoding

Papers read for the optimizations

Project Structure

.
├── backend/                 # Python FastAPI server
│   ├── app/                # Main application code
│   │   ├── core/          # Core LLM functionality
│   │   ├── optimizations/ # LLM optimization techniques
│   │   └── api/           # API endpoints
│   ├── tests/             # Backend tests
│   └── Dockerfile         # Backend Docker configuration
├── frontend/               # TypeScript/React frontend
│   ├── src/               # Source code
│   │   ├── components/    # React components
│   │   ├── hooks/        # Custom React hooks
│   │   └── types/        # TypeScript type definitions
│   └── Dockerfile         # Frontend Docker configuration
└── docker-compose.yml     # Docker compose configuration

About

Minimal server for local LLMs equipped with several inference optimization techniques.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages