A lot of this flew over my head, as I'm not familiar with the technology. But what I took away is that one of the reasons token generation is slow is because they're batching to save costs?
I am very annoyed that I ask a question (long one, a few sentences) and. many times I have to switch tasks while the LLM is completing their response. Would definitely be willing to pay more for fast replies.
Yes, and a corollary is that the more users you have, the lower the costs.
I think in the future there will be lots of users like you that value speed. That pushes for running fast LLM's on a laptop which would be interesting.
A lot of this flew over my head, as I'm not familiar with the technology. But what I took away is that one of the reasons token generation is slow is because they're batching to save costs?
I am very annoyed that I ask a question (long one, a few sentences) and. many times I have to switch tasks while the LLM is completing their response. Would definitely be willing to pay more for fast replies.
Yes, and a corollary is that the more users you have, the lower the costs.
I think in the future there will be lots of users like you that value speed. That pushes for running fast LLM's on a laptop which would be interesting.