Discussion about this post

User's avatar
Pratilipi's avatar

in general how we decide the number of experts, is it just hyperparameter tuning config or a standard approach for it, also regarding training losses what kind of losses works best pointwise or pairwise for mtl learning?

Leo's avatar

thank you for sharing! For the approaches you mentioned in task balancing section, which one is the best? Never heard of these approaches before. How did you find these papers?

2 more comments...

No posts

Ready for more?