Accuracy and Limitations
The reliability of our methodology varies with text length, as Figure 4 shows. For texts under 100 words, accuracy is limited to about 67%. As text length increases, accuracy improves dramatically, reaching 96% for texts exceeding 1,000 words.
This relationship between accuracy and text length reflects the fact that longer texts provide more linguistic data to identify patterns. For professional applications dealing with specialized texts (typically exceeding 500 words), our methodology provides highly reliable authentication with accuracy exceeding 90%.
Our validation testing confirms that, when properly applied, the methodology reliably distinguishes between human-written and AI-generated content. The multidimensional approach maintains effectiveness even as AI technologies evolve, as it focuses on fundamental linguistic differences rather than model-specific patterns.
Key limitations include the following:
- Effectiveness is limited with very short texts (under 100 words).
- Extensive human editing of AI-generated text can obscure some AI patterns, though many linguistic markers persist even after moderate editing.
- Evolving AI technology may change some linguistic markers.
In addition, while our methodology works well across multiple languages, different languages require specific calibration for optimal detection. Currently, English and Polish versions are the most refined, with research underway to extend the methodology to additional languages.
Conclusion
This hybrid methodology for text authentication integrates computational analysis with linguistic expertise to provide reliable probability assessments of AI involvement in text creation. By examining multiple linguistic features simultaneously — from statistical and probabilistic parameters to specific word usage patterns — our approach identifies the subtle but consistent differences that distinguish human writing from AI-generated content.
The methodology, derived from comprehensive analysis of current research and validated through empirical testing, achieves high accuracy rates, particularly for specialized texts exceeding 500 words. As language models continue to evolve, a multidimensional approach to text authentication will remain crucial. While specific linguistic markers may change, the fundamental differences in how humans and AI systems generate text provide enduring signals that can be used for authentication.
For professionals dealing with specialized texts where authenticity matters, this methodology offers a reliable tool for verification, contributing to greater transparency and accountability in an era of rapidly advancing language technologies.