I'd like posting privileges back, please. I feel bad for notifying everyone in this thread for such a stupid joke. not quite bad enough not to make it, though. we'll see if i regret it later.
An experiment I could conceivably do: what if you take a frozen existing transformer embedding, and then train an MLP to predict the second token in the dataset from the first token? Then, you freeze both, and fine-tune the transformer with the MLP acting as the LM head.