这个检查器能够而且无法检测到什么

每个受伤害的检查员都有盲点。 大部分埋在了这些盲点。 这页是我们的, 完整地写, 因为一份无法审问的报告 是一个你无法信任的报告。

引擎如何运作

检查分两个阶段进行。 首先,检索:我们从您文本的每个区域抽样8-12个字序列, 选择了比普通打开者更能检索的稀有文字, 并且以精确引用的短语在网上搜索每个在网上的文本。 这产生了一组候选页面 。

其次,验证:我们下载每个候选页面,并根据页面的实际内容,将您文本中的每句词句与本地内容相匹配。用滚动单词窗口比较,检测到单词窗口重叠;用象征性重叠和编辑相似的门槛来关闭副句句。 搜索引擎排在页面排位上,本身从未被当作证据来对待 — — 只有经核实的文本比较计数, 且每匹配都包含报告中的信任度数字 。

在此之前, 您的文本会正常化: 统一编码外观字符( 愚弄棋徒的标准把戏) 会被折回其直线等同符, 所以交换的西里尔字母不会隐藏匹配符 。

何时我们会错过匹配( 假底片)?

  • 付费和订阅来源:学术期刊、日志后面的新闻档案。
  • 私人数据库 - 包括诸如Turnitin的学术文件档案。 没有公共工具能看到它们; 任何自由检查者暗示其他信息都是骗你的。
  • 离线来源:从未上过网的印刷书籍和文件。
  • 内容非常新鲜:在几分钟或几个小时前,
  • 重副句: 重写更改大多数单词都低于近距离阈值的文字。 检测想法( 而不是文字) 超出了任何文本匹配器 。

何时我们标出无害文本(假正数)

  • 正确引用的材料 - 引用应该与其来源相符。 请检查引用, 而不是亮点 。
  • 常用短语和锅炉板:种群表达式、法律公式、整个领域重复的方法说明。
  • 书目和参考清单,自然与他们引用的作品相符。
  • 简短地说,即事实判决,是偶然的。

这就是为什么报告显示与你的文本相匹配的资料来源案文,按刑期分列,并分别标明准确和接近的对应案文:工具发现重叠;一个人判断它的意思。

分数意味着什么

匹配百分比是您在与检索源对齐的句子中的单词比例。 测量波段 - 0-5% 看上去原样, 6-20% 的比对, 21 的重要比对 - 正在根据典型的发文而不是指控阈值校准审查指南。 4%的报告仍然可以包含一个完全复制的段落,值得修正; 25% 正确引用的材料的报告可以完全诚实。

我们还告诉你,你检查了多少来源, 实际数字,打印在报告上, 因为无法核实的覆盖是市场营销, 而不是准确性。

Results are indicative, not conclusive. We compare your text against publicly accessible web pages at the moment you run the check. We cannot detect matches in sources that are offline, paywalled, unindexed, or held in private databases (including academic submission archives). Common phrases, correctly quoted material, and coincidental wording can appear as matches. Use this report as a guide for review and citation - not as standalone proof that text was or was not plagiarized.

经常问到的问题

Two phases. First we sample distinctive word sequences from your text and search them on the live web as exact phrases. Then we download every candidate page and align every sentence of your text against the page’s real content locally - exact matches by rolling word-window comparison, near matches by token overlap and edit similarity. Search results alone are never trusted as matches; only verified text comparison counts.

When the source is not on the public web at check time: paywalled journals, offline books, private submission archives, pages published minutes ago that search engines have not indexed, or content behind logins. Heavily paraphrased text can also fall below the near-match threshold. No web-based checker escapes these limits; we would rather tell you than let a 0% mislead you.

Correctly quoted passages, common stock phrases, technical boilerplate, legal or religious formulae, and bibliographies all legitimately match their sources. The report separates exact from near matches and shows the source text beside yours so a human can make the call - the score is an instrument reading, not a judgment.

It is the share of your words that sit in sentences we matched to a retrieved source: matched-sentence words divided by total words. The gauge bands are 0-5% "looks original", 6-20% "some matches - review citations", 21%+ "significant matching". The bands are review guidance, not accusation thresholds.

The interface runs in 100+ languages, and the engine itself is multilingual: sentence segmentation handles Latin, CJK and Arabic punctuation, and retrieval uses engines with strong non-English coverage. Match quality is best in languages with a large public web footprint; for very small languages, coverage is honestly thinner.