Of course you can generate text. You just need to run it in an autoregressive loop as a sort of reverse of all the fun Jev-like papers that have come out in the last couple days. Ask it to predict the next letter in a string, then sample at your favorite temperature, then predict the next letter, etc. This will be quite expensive, and it may work terribly. I’m not personally inclined to try it. I am, however, curious whether it would work less horribly if you correctly guess what tokenizer the input uses and request a choice over next tokens consistent with the tokenizer in question.
It would be absolutely hilarious if you did this, asked it which model it was, and it gave a recognizable answer that wasn’t Jev.
It would be absolutely hilarious if you did this, asked it which model it was, and it gave a recognizable answer that wasn’t Jev.