I’d go a step further and say your first point is a consequence of the second. Because “intelligence” is poorly defined, anyone can slap “AI” on their product, so it becomes a nebulous moving target.
Correct, we need new definitions and grading system to define what capabilities and algorithms has, while at the same time being on the lookout for new capabilities that fall outside of that system.
That way we can judge systems that are just a repeat of previous systems with dubious claims attached to their capabilities.
It can be somewhat difficult to do this properly and generally enough to ensure the tests aren't benchmaxxed.