AI workloads can be distributed across many processors.
Distrubuted Llama GitHub - b4rtaz/distributed-llama: Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
Exolabs https://exolabs.net/
Petals GitHub - bigscience-workshop/petals: 🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading https://medium.com/@visrow/decentralized-distributed-llm-using-ai-in-cost-effective-and-environment-friendly-way-8c0a73ee9e6f
AI Horde https://stablehorde.net/
Multiple GPUs can be used with llama.cpp now:
Merge into llama.cpp is not possible: