Next upHack for Humanity: San Francisco (powered by Google Gemini)
36:11

Reducing NLP Inference costs through model specialisation

This talk will discuss ways to reduce costs for NLP inference through a better choice of model, hardware, and model compression techniques.

Dmytro Spodarets
Dmytro Spodarets
Jul 10, 2023
Summary

Practical ways to reduce NLP inference costs through model specialisation, achieved with a better choice of model, the right hardware, and a range of model compression techniques applied together.