Interpreting how instruction-tuned transformers encode discourse relations, focusing on causation and antithesis, in next-token prediction tasks
Read the original at arxiv.org→arXiv:2607.18570v1 Announce Type: new Abstract: Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality. In this work, we investigate...
Original headline: "For What Reason? Interpreting Models' Encoding of Causation and Antithesis"