A new version of my book on cooperative multi-agent reinforcement learning is now available. A longer version will eventually be published with Frans Oliehoek so let us know your thoughts! arxiv.org/abs/2405.06161
- Amato organizes cooperative MARL around three settings: centralized training and execution, centralized training for decentralized execution, and decentralized training and execution.
- The paper walks through independent Q-learning, value factorization methods VDN, QMIX and QPLEX, and centralized critic methods MADDPG, COMA and MAPPO.
- Amato writes that CTDE is the most common paradigm because it leverages centralized information at training while keeping execution decentralized.