Семинары: Г. Малиновский, ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!

Семинары

RUS ENG

ЖУРНАЛЫ ПЕРСОНАЛИИ ОРГАНИЗАЦИИ КОНФЕРЕНЦИИ СЕМИНАРЫ ВИДЕОТЕКА ПАКЕТ AMSBIB

JavaScript is disabled in your browser. Please switch it on to enable full functionality of the website

	Календарь
	Поиск
	Регистрация семинара

	RSS
	Ближайшие семинары

Общероссийский семинар по оптимизации им. Б.Т. Поляка
20 апреля 2022 г. 17:30, Москва, Онлайн, пятница, 19:00

ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!

Г. Малиновский

Количество просмотров:
Эта страница:	238
Youtube:

https://youtu.be/xL-q3WCi_EU

Аннотация: We introduce ProxSkip – a surprisingly simple and provably efficient method for minimizing the sum of a smooth (f) and an expensive nonsmooth proximable (ψ) function. The canonical approach to solving such problems is via the proximal gradient descent (ProxGD) algorithm, which is based on the evaluation of the gradient of f and the prox operator of ψ in each iteration. In this work we are specifically interested in the regime in which the evaluation of prox is costly relative to the evaluation of the gradient, which is the case in many applications. ProxSkip allows for the expensive prox operator to be skipped in most iterations: while its iteration complexity is $\mathcal O(\kappa \log\frac{1}{\varepsilon})$, where κ is the condition number of f, the number of prox evaluations is $\mathcal O(\sqrt{\kappa} \log \frac{1}{\varepsilon})$ only. Our main motivation comes from federated learning, where evaluation of the gradient operator corresponds to taking a local GD step independently on all devices, and evaluation of prox corresponds to (expensive) communication in the form of gradient averaging. In this context, ProxSkip offers an effective acceleration of communication complexity. Unlike other local gradient-type methods, such as FedAvg, SCAFFOLD, S-Local-GD and FedLin, whose theoretical communication complexity is worse than, or at best matching, that of vanilla GD in the heterogeneous data regime, we obtain a provable and large improvement without any heterogeneity-bounding assumptions.

Обратная связь:

Пользовательское соглашение

Регистрация посетителей портала

Логотипы