Glossary

Inference Latency

تأخير الاستدلال

الوقت الإضافي المستغرق لتوليد المخرجات بسبب الحسابات المُضافة. LoRA لا يُضيف أي تأخير لأن مصفوفاته تُدمَج في الأوزان الأصلية.

Additional time required to generate outputs due to added computation. LoRA adds zero latency because its matrices merge into the original weights.

Also translated asزمن الاستجابة، وقت الاستنتاج

First appears in this corpus in: Fast Inference from Transformers via Speculative Decoding (2023)

Appears in these papers