Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The idea that human-readable explanations emitted by a language model don't necessarily correspond to the model's actual internal process of reaching a conclusion reminds me of parallel construction [1], a (fraudulent) law enforcement strategy of obtaining evidence of a crime through usually illegal means and claiming that the evidence was obtained legally through some other means.

[1] https://www.hrw.org/report/2018/01/09/dark-side/secret-origi...





Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: