Skills

All Skills

vllm

Skills tagged with #vllm

@vllm-project

vllm-deploy-docker

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills+1 more
5mo ago
450
@Orchestra-Research

awq-quantization

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.

OptimizationAWQQuantization4-BitActivation-AwareMemory Optimization
Orchestra-Research/AI-research-SKILLs+46 more
5mo ago
5.0K0
@vllm-project

main2main

The main2main skill guides an AI agent to adapt the latest vLLM main branch code for vLLM Ascend project.

vllm-project/vllm-ascend+2 more
5mo ago
1.8K0
@qso-graph
MCP

Io.Github.Qso Graph/Qsp Mcp

QSP — relay MCP tools to any OpenAI-compatible local LLM (llama.cpp, Ollama, vLLM)

mcpgithubaillm
qso-graph/qsp-mcp
5mo ago
0
@flagos-ai

kernelgen-flagos

Unified GPU kernel operator generation skill. Automatically detects the target repository type (FlagGems, vLLM, or general Python/Triton) and dispatches to the appropriate specialized sub-skill. Also includes a feedback submission sub-skill for bug reports. Use this skill when the user wants to generate a GPU kernel operator, create a Triton kernel, or says things like "generate an operator", "create a kernel for X", or "/kernelgen-flagos". This single skill replaces the need to install kernelgen-general, kernelgen-for-flaggems, kernelgen-for-vllm, and kernelgen-submit-feedback separately.

flagos-ai/KernelGen
5mo ago
300
@PrimeIntellect-ai

inference-server

Start and test the prime-rl inference server. Use when asked to run inference, start vLLM, test a model, or launch the inference server.

PrimeIntellect-ai/prime-rl+2 more
5mo ago
1.1K0