I know what the paper says but actually using both, GPT-4 is far ahead of google and Deepl. I think the isolated one sentence datasets used for evaluations are no longer up to snuff.
Trying something longer and more comprehensive and the difference is very clear.
Trying something longer and more comprehensive and the difference is very clear.
https://github.com/ogkalu2/Human-parity-on-machine-translati...