A. V. Kurtukova, A. S. Romanov, “Identification author of source code by machine learning methods”, Tr. SPIIRAN, 18:3 (2019), 742

Trudy SPIIRAN

RUS ENG

JOURNALS PEOPLE ORGANISATIONS CONFERENCES SEMINARS VIDEO LIBRARY PACKAGE AMSBIB

JavaScript is disabled in your browser. Please switch it on to enable full functionality of the website

	General information
	Latest issue
	Archive

	Search papers
	Search references

	RSS
	Latest issue
	Current issues
	Archive issues
	What is RSS

Informatics and Automation:
Year:
Volume:
Issue:
Page:
	Find

Personal entry:
Login:
Password:
	Save password
	Enter
	Forgotten password?
	Register

Trudy SPIIRAN, 2019, Issue 18, volume 3, Pages 742–766
DOI: https://doi.org/10.15622/sp.2019.18.3.741-765 (Mi trspy1062)

This article is cited in 5 scientific papers (total in 5 papers)

Artificial Intelligence, Knowledge and Data Engineering

Identification author of source code by machine learning methods

A. V. Kurtukova, A. S. Romanov

Tomsk State University of Control Systems and Radioelectronics (TUSUR)

Full-text PDF (1277 kB) Citations (5)

DOI: https://doi.org/10.15622/sp.2019.18.3.741-765

Abstract: The paper is devoted to the analysis of the problem of determining the source code author, which is of interest to researchers in the field of information security, computer forensics, assessment of the quality of the educational process, protection of intellectual property.
The paper presents a detailed analysis of modern solutions to the problem. The authors suggest two new identification techniques based on machine learning algorithms: support vector machine, fast correlation filter and informative features; the technique based on hybrid convolutional recurrent neural network.
The experimental database includes samples of source codes written in Java, C ++, Python, PHP, JavaScript, C, C # and Ruby. The data was obtained using a web service for hosting IT-projects – Github. The total number of source codes exceeds 150 thousand samples. The average length of each of them is 850 characters. The case size is 542 authors.
The experiments were conducted with source codes written in the most popular programming languages. Accuracy of the developed techniques for different numbers of authors was assessed using 10-fold cross-validation. An additional series of experiments was conducted with the number of authors from 2 to 50 for the most popular Java programming language. The graphs of the relationship between identification accuracy and case size are plotted. The analysis of result showed that the method based on hybrid neural network gives 97% accuracy, and it’s at the present time the best-known result. The technique based on the support vector machine made it possible to achieve 96% accuracy. The difference between the results of the hybrid neural network and the support vector machine was approximately 5%.

Keywords: source code writer, deep learning, neural network, SVM, HNN.

Received: 23.02.2019

Bibliographic databases:

Document Type: Article

UDC: 519.25: 004.8

Language: Russian

Citation: A. V. Kurtukova, A. S. Romanov, “Identification author of source code by machine learning methods”, Tr. SPIIRAN, 18:3 (2019), 742–766

Citation in format AMSBIB

\Bibitem{KurRom19}

\by A.~V.~Kurtukova, A.~S.~Romanov

\paper Identification author of source code by machine learning methods

\jour Tr. SPIIRAN

\yr 2019

\vol 18

\issue 3

\pages 742--766

\mathnet{http://mi.mathnet.ru/trspy1062}

\crossref{https://doi.org/10.15622/sp.2019.18.3.741-765}

\elib{https://elibrary.ru/item.asp?id=38515507}

Linking options:

https://www.mathnet.ru/eng/trspy1062

https://www.mathnet.ru/eng/trspy/v18/i3/p742

This publication is cited in the following 5 articles:

Citing articles in Google Scholar: Russian citations, English citations
Related articles in Google Scholar: Russian articles, English articles

Statistics & downloads:
Abstract page:	318
Full-text PDF :	219

Что такое QR-код?

Registration to the website

Logotypes