<?xml version="1.0" encoding="UTF-8"?>
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="JATS-archive-oasis-article1-4.xsd" article-type="research-article" dtd-version="1.4" xml:lang="ru">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Журнал Современные наукоемкие технологии</journal-title>
      </journal-title-group>
      <issn>1812-7320</issn>
      <publisher>
        <publisher-name>Общество с ограниченной ответственностью &amp;quot;Издательский Дом &amp;quot;Академия Естествознания&amp;quot;</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.17513/snt.40399</article-id>
      <article-id pub-id-type="publisher-id">ART-40399</article-id>
      <title-group>
        <article-title>ИНСТРУМЕНТЫ АВТОМАТИЗАЦИИ ДЛЯ ОБЕСПЕЧЕНИЯ ВОСПРОИЗВОДИМОСТИ ИССЛЕДОВАНИЙ В НАУКЕ О ДАННЫХ</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name-alternatives>
            <name xml:lang="ru">
              <surname>Горбунов</surname>
              <given-names>В.И.</given-names>
            </name>
          </name-alternatives>
          <name-alternatives>
            <name xml:lang="en">
              <surname>Gorbunov</surname>
              <given-names>V.I.</given-names>
            </name>
          </name-alternatives>
          <email>gorbunov.v93@gmail.com</email>
          <xref ref-type="aff" rid="aff1"/>
          <xref ref-type="aff" rid="aff2"/>
        </contrib>
        <contrib contrib-type="author">
          <name-alternatives>
            <name xml:lang="ru">
              <surname>Салимов</surname>
              <given-names>Т.А.</given-names>
            </name>
          </name-alternatives>
          <name-alternatives>
            <name xml:lang="en">
              <surname>Salimov</surname>
              <given-names>T.A.</given-names>
            </name>
          </name-alternatives>
          <email>ptrdiff.t@gmail.com</email>
          <xref ref-type="aff" rid="aff1"/>
        </contrib>
        <contrib contrib-type="author">
          <name-alternatives>
            <name xml:lang="ru">
              <surname>Горбань</surname>
              <given-names>Е.В.</given-names>
            </name>
          </name-alternatives>
          <name-alternatives>
            <name xml:lang="en">
              <surname>Gorban</surname>
              <given-names>E.V.</given-names>
            </name>
          </name-alternatives>
          <email>egor_gorbann@mail.ru</email>
          <xref ref-type="aff" rid="aff1"/>
        </contrib>
      </contrib-group>
      <aff id="aff1">
        <institution xml:lang="ru">ФГАОУ ВО «Национальный исследовательский институт ИТМО»</institution>
        <institution xml:lang="en">ITMO University</institution>
      </aff>
      <aff id="aff2">
        <institution xml:lang="ru">ФГБОУ ВО «Санкт-Петербургский государственный университет»</institution>
        <institution xml:lang="en">St. Petersburg State University</institution>
      </aff>
      <pub-date date-type="pub" iso-8601-date="2025-05-06">
        <day>06</day>
        <month>05</month>
        <year>2025</year>
      </pub-date>
      <issue>5</issue>
      <fpage>119</fpage>
      <lpage>126</lpage>
      <permissions>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>This is an open-access article distributed under the terms of the CC BY 4.0 license.</license-p>
        </license>
      </permissions>
      <self-uri content-type="url" hreflang="ru">https://top-technologies.ru/article/view?id=40399</self-uri>
      <abstract xml:lang="ru" lang-variant="original" lang-source="author">
        <p>Современные междисциплинарные проекты в области науки о данных характеризуются высокой сложностью, множеством участников и необходимостью координации организационных и технических процессов. Одной из ключевых проблем в таких проектах является обеспечение воспроизводимости методов и результатов исследований. Целью работы является проведение обзора современных практик и инструментов, направленных на повышение воспроизводимости в проектах науки о данных, и их анализ с точки зрения управления исследовательским процессом. Был проведен систематический обзор из более чем 50 публикаций за 2015-2025 годы, направленный на выявление современных организационных практик и инструментов, применяемых для обеспечения воспроизводимости в проектах науки о данных из научных и прикладных публикаций, документации инструментов в науке о данных и открытых репозиториев. Из них 30 работ легли в основу данного обзора. В работе рассмотрены пять ключевых категорий решений: контроль версий кода, данных и отчетов; управление зависимостями и средами исполнения; автоматизация процессов и оркестрация пайплайнов; стандартизация хранения данных; документирование и обеспечение прозрачности. Особое внимание уделено управленческому эффекту от их применения – снижению издержек, рисков и трудозатрат на коммуникации и выполнение типовых работ. Основными ограничениями внедрения инструментов воспроизводимости в организационные процессы остаются необходимость зрелой инфраструктуры, организационных изменений и обучения персонала. Представленные выводы могут быть использованы при разработке стандартов управления исследовательскими проектами, формировании корпоративной культуры прозрачности и выборе инструментов для применения.</p>
      </abstract>
      <abstract xml:lang="en" lang-variant="translation" lang-source="translator">
        <p>Modern interdisciplinary data science projects are characterized by high complexity, multiple participants, and the need to coordinate organizational and technical processes. One of the key challenges in such projects is to ensure reproducibility of research methods and results. The aim of the work is to review modern practices and tools aimed at improving reproducibility in data science projects, and to analyze them from the point of view of managing the research process. A systematic review of more than 50 publications from 2015-2025 was conducted, aimed at identifying modern organizational practices and tools used to ensure reproducibility in data science projects from scientific and applied publications, documentation of tools in data science and open repositories. Of these, 30 papers formed the basis of this review. The paper considers five key categories of solutions: version control of code, data and reports; dependency management and execution environments; automation of processes and pipeline orchestration; standardization of data storage; documentation and transparency. Special attention is paid to the management effect of their use – reducing costs, risks and labor costs for communication and standard work. The main limitations of implementing reproducibility tools in organizational processes remain the need for mature infrastructure, organizational changes, and staff training. The presented conclusions can be used in the development of standards for the management of research projects, the formation of a corporate culture of transparency and the selection of tools for application.</p>
      </abstract>
      <kwd-group xml:lang="ru">
        <kwd>управление исследованиями</kwd>
        <kwd>наука о данных</kwd>
        <kwd>воспроизводимость</kwd>
        <kwd>управление автоматизацией исследований</kwd>
        <kwd>организационно-технические системы</kwd>
      </kwd-group>
      <kwd-group xml:lang="en">
        <kwd>research management</kwd>
        <kwd>data science</kwd>
        <kwd>reproducibility</kwd>
        <kwd>research automation management</kwd>
        <kwd>socio-technical systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <back>
    <ref-list>
      <ref>
        <note>
          <p>1. Goodman S.N., Fanelli D., Ioannidis J.P.A. What does research reproducibility mean? // Science Translational Medicine. 2016. Vol. 8. P. 341. DOI: 10.1126/scitranslmed.aaf5027.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>2. Baker M. 1,500 scientists lift the lid on reproducibility // Nature. 2016. Vol. 533. P. 452-454. URL: https://www.nature.com/articles/533452a (дата обращения: 26.02.2025). DOI: 10.1038/533452a.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>3. Munafò M.R., Nosek B.A., Bishop D.V.M., Button K.S., Chambers C.D., Percie du Sert N., Simonsohn U., Wagenmakers E.-J., Ware J.J. Ioannidis J.P.A. A Manifesto for Reproducible Science // Nature Human Behaviour. 2017. Vol. 1, Is. 1. URL: https://www.nature.com/articles/s41562-016-0021 (дата обращения: 26.02.2025). DOI: 10.1038/s41562-016-0021.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>4. Stodden V. The data science life cycle: a disciplined approach to advancing data science as a science // Communications of the ACM. 2020. Vol. 63, Is. 7. P. 58-66. DOI: 10.1145/3360646.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>5. Gundersen O., Kjensmo S. State of the Art: Reproducibility in Artificial Intelligence // Proceedings – AAAI 2018 at New Orleans. Thirty-Second AAAI Conference on Artificial Intelligence 2018. 2018. URL: https://aaai.org/papers/11503-state-of-the-art-reproducibility-in-artificial-intelligence/ (дата обращения: 26.02.2025). DOI: 10.1609/aaai.v32i1.11503.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>6. Gundersen O.E., Coakley K., Kirkpatrick C., Gil Y. Sources of Irreproducibility in Machine Learning: A Review // arXiv:2204.07610v2[cs.LG]. 2023. DOI: 10.48550/arXiv.2204.07610.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>7. Waller L.A., Miller G.W. More than Manuscripts: Reproducibility, Rigor, and Research Productivity in the Big Data Era // Toxicological Sciences. 2016. Vol. 149, Is. 2. P. 275–276. URL: https://academic.oup.com/toxsci/article-abstract/149/2/275/2461691 (дата обращения: 26.02.2025). DOI: 10.1093/toxsci/kfv330.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>8. Liu J., Carlson J., Pasek J., Puchala B., Rao A., Jagadish H.V. Promoting and Enabling Reproducible Data Science Through a Reproducibility Challenge // Harvard Data Science Review. 2022. Vol. 4, Is. 3. URL: https://hdsr.mitpress.mit.edu/pub/mlconlea (дата обращения: 26.02.2025). DOI: 10.1162/99608f92.9624ea51.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>9. Hernandez J.A., Colom M. Repeatability, Reproducibility, Replicability, Reusability (4R) in Journals’ Policies and Software/Data Management in Scientific Publications: A Survey, Discussion, and Perspectives // arXiv preprint arXiv:2312.11028v1. 2023. DOI: 10.48550/arXiv.2312.11028.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>10. Chen K.Y., Toro-Moreno M., Subramaniam A.R. GitHub is an effective platform for collaborative and reproducible laboratory research // arXiv preprint arXiv:2408.09344v2. 2025. URL: https://arxiv.org/abs/2408.09344v2 (дата обращения: 26.02.2025).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>11. Idowu S., Osman O., Strüber D., Berger T. Machine learning experiment management tools: a mixed-methods empirical study // Empirical Software Engineering. 2024. Vol. 29, Is.4. DOI: 10.1007/s10664-024-10444-w.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>12. Klump J., Wyborn L., Wu M., Martin J., Downs R.R., Asmi A. Versioning data is about more than revisions: A conceptual framework and proposed principles // Data Science Journal. 2021. Vol. 20, Is. 1. P. 12. DOI: 10.5334/dsj-2021-012.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>13. Armbrust M., Das T., Sun L., Yavuz B., Zhu S., Murthy M., Torres J., van Hovell H., Ionescu A., Łuszczak A., Świtakowski M., Szafrański M., Li X., Ueshin T., Mokhtar M., Boncz P., Ghodsi A., Paranjpye S., Senster P., Xin R., Zaharia M. Delta Lake: high-performance ACID table storage over cloud object stores // Proc VLDB Endow. 2020. Vol. 13, Is. 12. P. 3411-3424. DOI: 10.14778/3415478.3415560.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>14. Semmelrock H., Ross-Hellauer T., Kopeinik S., Theiler D., Haberl A., Thalmann S., Kowald D. Reproducibility in Machine Learning-based Research: Overview, Barriers and Drivers // arXiv preprint arXiv:2406.14325. 2024. DOI: 10.48550/arXiv.2406.14325.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>15. Arpteg A., Brinne B., Crnkovic-Friis L., Bosch J. Software Engineering Challenges of Deep Learning // 2018 44th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), Prague, Czech Republic. 2018. P. 50-59. URL: https://ieeexplore.ieee.org/document/8498185 (дата обращения: 26.02.2025). DOI: 10.1109/SEAA.2018.00018.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>16. Nüst D., Sochat V., Marwick B., Eglen S.J., Head T., Hirst T., Evans B.D. Ten simple rules for writing Dockerfiles for reproducible data science // PLOS Computational Biology. 2020. Vol. 16, Is. 11. P. 1-24. DOI: 10.1371/journal.pcbi.1008316.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>17. Chirigati F., Rampin R., Shasha D., Freire J. ReproZip: Computational Reproducibility with Ease // Proceedings of the 2016 ACM SIGMOD International Conference on Management of Data, ACM. 2016. P. 2085-2088. DOI: 10.1145/2882903.2899401.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>18. Martínez J. de la R., Buso F., Kouzoupis A., Ormenisan A.A., Niazi S., Bzhalava D., Mak K., Jouffrey V., Ronström M., Cunningham R., Zangis R., Mukhedkar D., Khazanchi A., Vlassov V., Dowling J. The Hopsworks Feature Store for Machine Learning // SIGMOD/PODS ‘24 – Companion of the 2024 International Conference on Management of Data, Santiago AA, Chile. 2024. P. 135-147. DOI: 10.1145/3626246.3653389.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>19. Sculley D., Holt G., Golovin D., Davydov E., Phillips T., Ebner D., Chaudhary V., Young M., Crespo J.-F., Dennison D. Hidden Technical Debt in Machine Learning Systems // Proceedings of the 29th International Conference on Neural Information Processing Systems 2015. Vol. 2. P. 2503-2511. DOI: 10.5555/2969442.2969519.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>20. Sculley D., Holt G., Golovin D., Davydov E., Phillips T., Ebner D., Chaudhary V., Young M. Machine Learning: The High Interest Credit Card of Technical Debt // SE4ML: Software Engineering for Machine Learning (NIPS 2014 Workshop). 2014. URL: https://research.google/pubs/machine-learning-the-high-interest-credit-card-of-technical-debt/ (дата обращения: 26.02.2025).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>21. Di Tommaso P., Chatzou M., Floden E.W., Barja P.P., Palumbo E., Notredame C. Nextflow enables reproducible computational workflows // Nature Biotechnology. 2017. Vol. 35, Is. 4. P. 316-319. URL: https://www.nature.com/articles/nbt.3820 (дата обращения: 26.02.2025). DOI: 10.1038/nbt.3820.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>22. Zaharia M.A., Chen A., Davidson A., Ghodsi A., Hong S.A., Konwinski A., Murching S., Nykodym T., Ogilvie P., Parkhe M., Xie F., Zumar C. Accelerating the Machine Learning Lifecycle with MLflow // IEEE Data Eng Bull. 2018. Vol. 41, P. 39-45. URL: https://api.semanticscholar.org/CorpusID:83459546 (дата обращения: 26.02.2025).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>23. Yasmin J., Wang J., Tian Y., Adams B. An Empirical Study of Developers’ Challenges in Implementing Workflows as Code: A Case Study on Apache Airflow // Journal of Systems and Software. 2024. Vol. 219. DOI: 10.48550/arXiv.2406.00180.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>24. Freire J., Koop D., Santos E., Silva C. Provenance for Computational Tasks: A Survey // Computing in Science and Engineering. 2008. V. 10, Is. 3, P. 11-21. URL: https://ieeexplore.ieee.org/document/4488060 (дата обращения: 26.02.2025). DOI: 10.1109/MCSE.2008.79.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>25. Subramaniam P., Ma Y., Li C., Mohanty I., Fernandez R.C. Comprehensive and comprehensible data catalogs: The what, who, where, when, why, and how of metadata management // arXiv preprint arXiv:2103.07532. 2021. DOI: 10.48550/arXiv.2103.07532.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>26. Wilson G., Bryan J., Cranston K., Kitzes J., Nederbragt L., Teal T.K. Good enough practices in scientific computing // PLOS Computational Biology. 2017. Vol.13, Is.6. P.1-20. DOI: 10.1371/journal.pcbi.1005510.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>27. Pimentel J.F., Murta L., Braganholo L., Freire J. A Large-Scale Study About Quality and Reproducibility of Jupyter Notebooks // Proceedings – 2019 IEEE/ACM 16th International Conference on Mining Software Repositories, Montreal, QC, Canada. 2019. P. 507-517. URL: https://ieeexplore.ieee.org/document/8816763 (дата обращения: 26.02.2025). DOI: 10.1109/MSR.2019.00077.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>28. Amershi S., Begel A., Bird C., DeLine R., Gall H., Kamar E., Nagappan N., Nushi B., Zimmermann T. Software Engineering for Machine Learning: A Case Study // IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), Montreal, QC, Canada. 2019. P. 291-300. URL: https://ieeexplore.ieee.org/document/8804457 (дата обращения: 26.02.2025). DOI: 10.1109/ICSE-SEIP.2019.00042.</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>29. Salama Kh., Kazmierczak, J., Schut D. Practitioners Guide to MLOps: A Framework for Continuous Delivery and Automation of Machine Learning // Google Cloud White Paper. 2021. 37 р. URL: https://services.google.com/fh/files/misc/practitioners_guide_to_mlops_whitepaper.pdf (дата обращения: 26.02.2025).</p>
        </note>
      </ref>
      <ref>
        <note>
          <p>30. Treveil M., Omont N., Stenac C., Lefevre K., Phan D., Zentici J., Lavoillotte A., Miyazaki M., Heidmann L. Introducing MLOps: How to Scale Machine Learning in the Enterprise. O’Reilly Media, 2020. 183 р.</p>
        </note>
      </ref>
    </ref-list>
  </back>
</article>
