AI Performance

Distributed

AI workloads can be distributed across many processors.

Distribution

Parallel

Multiple GPUs can be used with llama.cpp now:

llama.cpp

ik_llama

Merge into llama.cpp is not possible:

Swarm

OptiLLM