<?xml version="1.0" encoding="UTF-8"?>
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="JATS-archive-oasis-article1-4.xsd" article-type="research-article" dtd-version="1.4" xml:lang="ru">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Журнал Современные наукоемкие технологии</journal-title>
      </journal-title-group>
      <issn>1812-7320</issn>
      <publisher>
        <publisher-name>Общество с ограниченной ответственностью &amp;quot;Издательский Дом &amp;quot;Академия Естествознания&amp;quot;</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.17513/snt.38954</article-id>
      <article-id pub-id-type="publisher-id">ART-38954</article-id>
      <title-group>
        <article-title>Векторное представление слов русского языка посредством нейросетевых моделей сверточного автоэнкодера</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name-alternatives>
            <name xml:lang="ru">
              <surname>Лихачев</surname>
              <given-names>А.Ю.</given-names>
            </name>
          </name-alternatives>
          <name-alternatives>
            <name xml:lang="en">
              <surname>Likhachev</surname>
              <given-names>A.Y.</given-names>
            </name>
          </name-alternatives>
          <email>andreyyy.lihachev@gmail.com</email>
          <xref ref-type="aff" rid="aff1"/>
        </contrib>
        <contrib contrib-type="author">
          <name-alternatives>
            <name xml:lang="ru">
              <surname>Трубянов</surname>
              <given-names>А.Б.</given-names>
            </name>
          </name-alternatives>
          <name-alternatives>
            <name xml:lang="en">
              <surname>Trubyanov</surname>
              <given-names>A.B.</given-names>
            </name>
          </name-alternatives>
          <email>true47@mail.ru</email>
          <xref ref-type="aff" rid="aff1"/>
        </contrib>
      </contrib-group>
      <aff id="aff1">
        <institution xml:lang="ru">ФГБОУ ВО «Марийский государственный университет»</institution>
        <institution xml:lang="en">Mari State University</institution>
      </aff>
      <pub-date date-type="pub" iso-8601-date="2021-12-27">
        <day>27</day>
        <month>12</month>
        <year>2021</year>
      </pub-date>
      <issue>12</issue>
      <fpage>52</fpage>
      <lpage>59</lpage>
      <permissions>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>This is an open-access article distributed under the terms of the CC BY 4.0 license.</license-p>
        </license>
      </permissions>
      <self-uri content-type="url" hreflang="ru">https://top-technologies.ru/article/view?id=38954</self-uri>
      <abstract xml:lang="ru" lang-variant="original" lang-source="author">
        <p>Исследование было проведено при использовании моделей нейронных сетей на основе аналитических методов и проведения эксперимента. При построении моделей использовался язык программирования Python и фреймворк машинного обучения Keras, встроенный в TensorFlow (открытая программная библиотека для машинного обучения, разработанная компанией Google для решения задач построения и тренировки нейронных сетей). Результаты работы представляют собой модель сверточного автокодировщика, состоящего из энкодера и декодера. Представленная модель оптимизирована относительно архитектуры нейронной сети и набора гиперпараметров с точки зрения минимизации длины вектора представления и минимизации ошибок восстановления слов после расшифровки (декодера). Разработанная модель представлена в виде базы для обучения, состоящей из 6,807724 млн слов русского языка, программной реализации на языке Python, гиперпараметров модели, набора параметров обученной сети. Обученная модель, примененная к обучающей выборке, восстанавливает (расшифровывает) слова при уровне ошибок – 0 %, примененная к тестовой выборке – 0,02 %. Построенная модель может быть использована для построения алгоритмов обработки естественного языка, в том числе для разработки интеллектуального дистанционного электронного помощника по взаимодействию информационных систем с пользователем.</p>
      </abstract>
      <abstract xml:lang="en" lang-variant="translation" lang-source="translator">
        <p>The study was conducted with the use of neural network models, on the basis of analytical methods and experimentation. The models were constructed with the use of the Python programming language and the Keras machine learning framework embedded in TensorFlow (an open software library for machine learning). The results present a convolutional autoencoder model, which consists of an encoder and a decoder. The presented model is optimized with respect to neural network architecture and the set of hyperparameters in terms of minimising the length of word embedding and minimising errors in words recovery after decoding. The developed model is presented as a training base consisting of 6.807724 million words of the Russian language, software implementation in the Python language, model hyperparameters and the set of parameters of the trained network. The trained model applied to the training set recovers (decrypts) words at the error level of 0 %, and the model applied to the test set recovers at the error level of 0.02 %. The constructed model can be used to develop algorithms of Natural Language Processing, including the development of remote intelligent digital assistant for interaction between information systems and a user.</p>
      </abstract>
      <kwd-group xml:lang="ru">
        <kwd>искусственный интеллект</kwd>
        <kwd>обработка естественного языка</kwd>
        <kwd>русский язык</kwd>
        <kwd>сверточные нейронные сети</kwd>
        <kwd>автокодировщик</kwd>
      </kwd-group>
      <kwd-group xml:lang="en">
        <kwd>artificial intelligence</kwd>
        <kwd>natural language processing</kwd>
        <kwd>Russian language</kwd>
        <kwd>convolutional neural networks</kwd>
        <kwd>autoencoder</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <back>
    <ref-list>
      <ref>
        <note>
          <p>1. Bullinaria J.A., Levy J.P. Extracting semantic representations from word co-occurrence statistics: A computational study. Behavior Research Methods. 2007. Vol. 39. No. 3. P. 510–526.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>2. Baroni M., Lenci A. Distributional memory: A general framework for corpus-based semantics. Computational Linguistics. 2010. Vol. 36. No. 4. P. 673–721.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>3. Cohen R., Goldberg Y., Elhadad M. Domain adaptation of a dependency parser with a class-class selectional preference model. In Proceedings of ACL 2012 Student Research Workshop (ACL ‘12). Association for Computational Linguistics. 2012. P. 43–48.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>4. Mikolov T., Kombrink S., Burget L., Cernocky J.H., Khudanpur S. Extensions of recurrent neural network language model. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2011. P. 5528–5531.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>5. Mikolov T., Sutskever I., Chen K., Corrado G.S., Dean J. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5–8, 2013, Lake Tahoe, Nevada, United States, 2013. P. 3111–3119.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>6. Collobert R., Weston J., Bottou L., Karlen M., Kavukcuoglu K., Kuksa P. Natural language processing (almost) from scratch. The Journal of Machine Learning Research. 2011. Vol. 12. P. 2493–2537.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>7. Socher R., Pennington J., Huang E.H., Ng A.Y., Manning C.D. Semi-supervised recursive autoencoders for predicting sentiment distributions. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, USA, 2011. P. 151–161.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>8. Al-Rfou R., Perozzi B., Skiena S. Polyglot: Distributed word representations for multilingual NLP. In Proc. of CoNLL 2013. P. 183–192.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>9. Krizhevsky A., Sutskever I., Hinton G.E. ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems 25 (NIPS 2012). [Electronic resource]. URL: http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks (date of access: 27.11.2021).</p>
        </note>
      </ref>
    </ref-list>
  </back>
</article>
