<?xml version="1.0" encoding="UTF-8"?>
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="JATS-archive-oasis-article1-4.xsd" article-type="research-article" dtd-version="1.4" xml:lang="ru">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Журнал Современные наукоемкие технологии</journal-title>
      </journal-title-group>
      <issn>1812-7320</issn>
      <publisher>
        <publisher-name>Общество с ограниченной ответственностью &amp;quot;Издательский Дом &amp;quot;Академия Естествознания&amp;quot;</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">ART-36585</article-id>
      <title-group>
        <article-title>ПРОГРАММЫ-КРАУЛЕРЫ ДЛЯ СБОРА ДАННЫХ О ПРЕДСТАВИТЕЛЬСКИХ САЙТАХ ЗАДАННОЙ ПРЕДМЕТНОЙ ОБЛАСТИ – АНАЛИТИЧЕСКИЙ ОБЗОР</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name-alternatives>
            <name xml:lang="ru">
              <surname>Печников</surname>
              <given-names>А.А.</given-names>
            </name>
          </name-alternatives>
          <name-alternatives>
            <name xml:lang="en">
              <surname>Pechnikov</surname>
              <given-names>A.A.</given-names>
            </name>
          </name-alternatives>
          <email>pechnikov@krc.karelia.ru</email>
          <xref ref-type="aff" rid="aff1"/>
        </contrib>
        <contrib contrib-type="author">
          <name-alternatives>
            <name xml:lang="ru">
              <surname>Сотенко</surname>
              <given-names>Е.М.</given-names>
            </name>
          </name-alternatives>
          <name-alternatives>
            <name xml:lang="en">
              <surname>Sotenko</surname>
              <given-names>E.M.</given-names>
            </name>
          </name-alternatives>
          <email>katherinmail@gmail.com</email>
          <xref ref-type="aff" rid="aff2"/>
        </contrib>
      </contrib-group>
      <aff id="aff1">
        <institution xml:lang="ru">ФГБУН «Институт прикладных математических исследований Карельского научного центра Российской академии наук»</institution>
        <institution xml:lang="en">Institute of Applied Mathematical Research of the Karelian Research Centre of the Russian Academy of Sciences</institution>
      </aff>
      <aff id="aff2">
        <institution xml:lang="ru">ФГБОУ ВО «Санкт-Петербургский государственный университет»</institution>
        <institution xml:lang="en">St. Petersburg State University</institution>
      </aff>
      <pub-date date-type="pub" iso-8601-date="2017-02-21">
        <day>21</day>
        <month>02</month>
        <year>2017</year>
      </pub-date>
      <issue>2</issue>
      <fpage>58</fpage>
      <lpage>62</lpage>
      <permissions>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>This is an open-access article distributed under the terms of the CC BY 4.0 license.</license-p>
        </license>
      </permissions>
      <self-uri content-type="url" hreflang="ru">https://top-technologies.ru/article/view?id=36585</self-uri>
      <abstract xml:lang="ru" lang-variant="original" lang-source="author">
        <p>Анализ данных, представленных в Вебе, – распространенная на сегодняшний день исследовательская задача. Для её решения необходимы программы-инструменты, позволяющие собирать данные из Веба. Для обозначения таких программ часто используется термин «краулер». Сегодня существует широкий спектр известных краулеров, и в ряде случаев это позволяет не писать с нуля новые, а использовать существующие программы. Если это краулеры с открытым кодом, то их можно дорабатывать под цели конкретных задач, формирующиеся и видоизменяющиеся в процессе исследования. В данной работе в качестве альтернативы рассмотрены несколько краулеров, из которых требуется сделать выбор наиболее подходящего для решения задачи сбора данных с некоторого заданного ограниченного множества представительских сайтов (таких, например, как сайты гостиниц). Приводятся основные требования, предъявляемые к краулерам, и их классификация по основным типам. Дается аналитический обзор наиболее популярных краулеров, для которых в качестве одного из главных критериев отбора является наличие открытого исходного кода. В результате проведенного анализа проведен отбор трех наиболее перспективных краулеров.</p>
      </abstract>
      <abstract xml:lang="en" lang-variant="translation" lang-source="translator">
        <p>The analysis of data, presented in Web, widely spread nowadays research task. In order to solve it usually used programs, gathering data from the Web, which are called «crawlers». There are a lot of already built crawlers, that allows not to write your own, but use one of existing instead. Some crawlers are open source that allows to modify them in case they don’t provide required functionality. This paper provides the overview of several crawlers in order to solve the task of gathering the data from set of official sites (for example, official sites of hotels). Also were provided main requirements, which crawlers must meet, crawler’s classification and analysis of their functionality. Is the crawler open source – was one of the main criteria of crawler’s selection. The result of performed analysis was the selection of three most perspective crawlers.</p>
      </abstract>
      <kwd-group xml:lang="ru">
        <kwd>Интернет</kwd>
        <kwd>Веб</kwd>
        <kwd>сбор данных о Вебе</kwd>
        <kwd>краулер</kwd>
        <kwd>открытое программное обеспечение</kwd>
      </kwd-group>
      <kwd-group xml:lang="en">
        <kwd>Internet</kwd>
        <kwd>Web</kwd>
        <kwd>web data collection</kwd>
        <kwd>crawler</kwd>
        <kwd>open source</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <back>
    <ref-list>
      <ref>
        <note>
          <p>1. Apache Cassandra. http://cassandra.apache.org (дата обращения: 06.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>2. Apache Lucene. http://lucene.apache.org (дата обращения: 06.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>3. Apache Nutch™. http://nutch.apache.org (дата обращения: 01.03.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>4. Apache Hadoop. https://wiki.apache.org/hadoop (дата обращения: 06.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>5. Bar-Ilan J. Data collection methods on the Web for infometric purposes: A review and analysis // Scientometrics. January 2001. – Vol. 50(1). – Р. 7–32.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>6. Best Open Source: 36 best open source web crawler projects. http://www.findbestopensource.com/tagged/webcrawler start=0 (дата обращения: 11.01.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>7. Bixo – Web Mining Toolkit. https://openbixo.files.wordpress.com/2010/01/bixo-web-mining-talk-at-hug.pdf (дата обращения: 02.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>8. Cascading | Application Platform for Enterprise Big Data. http://www.cascading.org (дата обращения: 06.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>9. Data Warehousing Overview. http://dwreview.com/DW_Overview.html (дата обращения: 06.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>10. Examples using Common Crawl Data. http://commoncrawl.org/the-data/examples (дата обращения: 06.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>11. Granados N.F., Kauffman R.J., King B. The Emerging role of vertical search engines in travel distribution: a newly-vulnerable electronic markets perspective // Proceedings of the 41st Hawaii International Conference on System Sciences. – 2008. – P. 389.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>12. GitHub – yasserg/crawler4j: Open Source Web Crawler for Java. https://github.com/yasserg/crawler4j (дата обращения: 01.03.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>13. Home – arachnode.net. http://arachnode.net (дата обращения: 01.03.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>14. Khurana D., Kumar S. Web Crawler: A Review // International Journal of Computer Science &amp; Management Studies. – 2012. – № 12–1. – P. 401–405.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>15. Leading Enterprise Java Web Framework | ZK. https://www.zkoss.org (дата обращения: 06.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>16. Norconex HTTP Collector. http://www.norconex.com/product/collector-http (дата обращения: 01.03.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>17. Olston C., Najork M. Web Crawling // Foundations and Trends in Information Retrieval. – 2010. – Vol. 4, № 3. – P. 175–246.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>18. OpenSearchServer | Open Source Search Engine and Search API. http://www.open-search-server.com (дата обращения: 01.03.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>19. Papavassiliou V., Prokopidis V., Thurmair G. A modular open-source focused crawler for mining monolingual and bilingual corpora from the web // Proceedings of the 6th Workshop on Building and Using Comparable Corpora., Sofia, August 8, 2013. – P. 43–51.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>20. Performance Benchmark. https://arachnode.net/blogs/arachnode_net/archive/2008/12/31/performance-benchmark-etc.aspx (дата обращения: 4.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>21. Performance Tuning. https://arachnode.net/blogs/arachnode_net/archive/2015/04/12/performance-tuning.aspx (дата обращения: 04.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>22. RFC 3986 – Uniform Resource Identifier (URI). 2005. https://tools.ietf.org/html/rfc3986 (дата обращения: 04.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>23. Ricardo A., Serrao C. Comparison of existing open-source tools for Web crawling and indexing of free Music // Journal of telecommunications. – 2013. – № 18–1. https://ru.scribd.com/doc/123153248/Comparison-of-existing-open-source-tools-for-Web-crawling-and-indexing-of-free-Music.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>24. Scrapinghub Platform. http://scrapinghub.com/platform (дата обращения: 12.01.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>25. Scrapy A Fast and Powerful Scraping and Web Crawling Framework. http://scrapy.org (дата обращения: 01.03.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>26. Sharma G., Sharma S., Singla H. Evolution of web crawler its challenges // International Journal of Computer Technology and Applications. – 2016. – № 9(11). – P. 5357–5368.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>27. Top 50 open source web crawlers for data mining. http://bigdata-madesimple.com/top-50-open-source-web-crawlers-for-data-mining (дата обращения: 11.01.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>28. Udapure T.V., Kale R.D., Dharmik R.C. Study of Web Crawler and its Different Types // Journal of Computer Engineering. 2014. № 16-1. http://iosrjournals.org/iosr-jce/papers/Vol16-issue1/Version-6/A016160105.pdf.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>29. Web Crawling with Apache Nutch. http://events.linuxfoundation.org/sites/events/files/slides/aceu2014-snagel-web-crawling-nutch.pdf (дата обращения: 05.02.2017).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>30. Wget – GNU Project – Free Software Foundation. http://www.gnu.org/software/wget (дата обращения: 01.03.2017).</p>
        </note>
      </ref>
    </ref-list>
  </back>
</article>
