Researchers reverse-engineer LLM prompts from output accurately
Researchers at IIT Bombay and Adobe Research demonstrate a method that reconstructs proprietary LLM prompts from model outputs with near-perfect accuracy without requiring access to model weights.
Researchers at IIT Bombay and Adobe Research have built an inverse language model that reconstructs the original prompt from an LLM's output with near-perfect accuracy. Their method, called "Previous-Token Prediction," doesn't need access to model weights and works across different models.
For companies relying on proprietary system prompts, this could be a serious security risk. The article Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy appeared first on The Decoder.