Distributed extra-gradient with optimal complexity and communication guarantees
February 1, 2023
When using quantization for distributed training of GANs or multi-agent RL, the additional variance due to quantization can slow down convergence. In this paper we formalize this, show that extra-gradient type methods with adaptive quantization can maintain the optimal oracle complexity while decreasing communication overhead, and show this theory translates to practice on standard benchmarks.