Visitar URL original
GitHub - Nicoveu/SDG_Semanticcode · GitHub
Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

 
 

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

FSME

FCHE (Fine-grained Semantic Metrics Extractor ) is a tool for extraction of file—level semantic metrics.

Semantic Similarity Matrix

Given a set of documents, D = D1, D2, . . . Dn, and their related vectors V = V1, V2, . . . , Vn. The semantic similarity between any two documents could be calculated by taking the cosine between their corresponding vectors. In this step, we generated a matrix to represent the semantic similarity between any two files. The similarity matrix is a square matrix (N × N matrix) that consists of a set of source files. Each cell in this matrix, cell(fi, fj), shows the semantic similarity between files i and j, that is, the semantic similarity of documents Di and Dj.

Semantic dependency graph

Based on the semantic similarity matrix, we further transferred it into a semantic dependency graph, G. G =< V, E >,where V is the set of source files and E is the set of edges.Each of the edges, e(fi, fj), indicates that file i semantically depends file j. The weight of each edge will be the value of corresponding semantic similarity. In this case, we consider there exists an edge from file A to B (i.e. file A semantically depends on file B), if the similarity weight between A and B satisfies: w(A, B) >= C, where C is a threshold indicting the significance of the similarity. If a value is larger than the threshold, we consider there is a semantic dependency.Following the well-know Pareto principle, we set C to be 0.8, meaning that when the w(A, B) >= 0.8, we consider the semantic dependency from file A to B is significant. In addition, to substantiate this setting, we randomly selected one hundred pairs of files and manually examined their lexicon information. We have observed that, when the similarity value of two files is larger than 0.8, we can explicitly identify similar or even identical names, identifiers, or comments. ect.

Metrics

Semantic Metric

Semantic Metric Description
Dependent Count (DC) Number of files that semantically depends on the file.
Sum of Dependent Weight (SDW) Sum of the semantic weight of the edges formed by fileand the files that semantically depend on it.
Average of Dependency Weight (ADW) Average of the semantic weight of the edges formed by file and the files that semantically depend on it: ADW=SDW/DC.
Longest Dependency Path(LDP) Number of files involved in the longest path of file in the semantic dependency graph.
Impact Space File Count (ISF) Number of files involved in file’s impact space.
Sum of Files’ Weight (SFW) Sum of similarity weights between file and the other files involved in its impact space.
Average of Files’ Weight (AFW) Average of similarity weights between file and the other files inits impact space .AFW=SFW/ISF
Impact Space Weight (ISW) Sum of the weights of all edges in the sub graph (i.e. the impact space).
Average Weight of Impact Space (AWIS) Average weight of file’s impact space: AWIS=ISWf/n, where n is the number of edges in the impact space.
Semantic Similarity Sum (SSS) Sum of the similarity between a file and all files in the same project.
Semantic similarity average (SSA) Average of the similarity between a file and all files in the same project.

Features

Usage

1) Set up Java environment and Python environment

you should set up Java environment(jdk1.8). you should set up Python3.0 environment(64 bit). you should set up understand environment(64 bit).

2)Use

before compile the source code,you should set Program arguments: < projectName_Write > < projectName > < readString > < writeString > < udb_path > < Write_path >

  • < projectName_Write >. The folder for storing data created by the project name, for example: projectName_Write + "_data".
  • < ** projectName** >. The name of the project.
  • < filename.csv >. Project file name exported with [understand]. (https://scitools.com/).
  • < readString >. Storage path of project file name.
  • < writeString >.Storage path of exported semantic indicators
  • < udb_path >. The path of the UDB generated by the project import [understand] (https://scitools.com/).
  • < Write_path >.The storage path of the intermediate file generated by the program.

Please be sure to use files in the same format as the example, and ensure that they are in the same order, otherwise the results may be inaccurate.

Example

before run this project, you should Compile Configurations and set Program arguments: projectName_Write = "xerces1.2" projectName = "xerces-1.2.0-src" readString = "C:\Users\sherlock\Apach\" + projectName_Write + "_date\" + projectName_Write + "_filename.csv" writeString = "C:\Users\sherlock\Apach\" + projectName_Write + "_data\" + projectName_Write + "_filename_metric.csv" udb_path = "C:\Users\sherlock\Apach\" + projectName + "\" + projectName + ".udb" Write_path="C:\Users\sherlock\Apach\" + projectName_Write + "_dataTime\" + projectName_Write

then run the project, you will get xerces1.2_filename_metric.csv in project.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages