维普中文期刊产品整合服务
376篇 您的检索式:期刊名="Data science journal"
    题名 作者 年代 出处 被引量
1Science Mapping:A Systematic Review of the Literature显示文摘Purpose: We present a systematic review of the literature concerning major aspects of science mapping to serve two primary purposes: First, to demonstrate the use of a science mapping approach to perform the review so that researchers may apply the procedure to the review of a scientific domain of their own interest, and second, to identify major areas of research activities concerning science mapping, intellectual milestones in the development of key specialties, evolutionary stages of major specialties involved, and the dynamics of transitions from one specialty to another.Design/methodology/approach: We first introduce a theoretical framework of the evolution of a scientific specialty. Then we demonstrate a generic search strategy that can be used to construct a representative dataset of bibliographic records of a domain of research. Next, progressively synthesized co-citation networks are constructed and visualized to aid visual analytic studies of the domain's structural and dynamic patterns and trends. Finally, trajectories of citations made by particular types of authors and articles are presented to illustrate the predictive potential of the analytic approach.Findings: The evolution of the science mapping research involves the development of a number of interrelated specialties. Four major specialties are discussed in detail in terms of four evolutionary stages: conceptualization, tool construction, application, and codification. Underlying connections between major specialties are also explored. The predictive analysis demonstrates citations trajectories of potentially transformative contributions.Research limitations: The systematic review is primarily guided by citation patterns in the dataset retrieved from the literature. The scope of the data is limited by the source of the retrieval, i.e. the Web of Science, and the composite query used. An iterative query refinement is possible if one would like to improve the data quality, although the current approach serves our purpose adequately. More in-depth analyses of each specialty would be more revealing by incorporating additional methods such as citation context analysis and studies of other aspects of scholarly publications.Practical implications: The underlying analytic process of science mapping serves many practical needs, notably bibliometric mapping, knowledge domain visualization, and visualization of scientific literature. In order to master such a complex process of science mapping, researchers often need to develop a diverse set of skills and knowledge that may span multiple disciplines. The approach demonstrated in this article provides a generic method for conducting a systematic review.Originality/value: Incorporating the evolutionary stages of a specialty into the visual analytic study of a research domain is innovative. It provides a systematic methodology for researchers to achieve a good understanding of how scientific fields evolve, to recognize potentially insightful patterns from visually encoded signs, and to synthesize various information so as to capture the state of the art of the domain.Chaomei Chen 2017Journal of Data and Information Science2017,2,2:596
2Masked Sentence Model Based on BERT for Move Recognition in Medical Scientific Abstracts显示文摘Purpose:Mo ve recognition in scientific abstracts is an NLP task of classifying sentences of the abstracts into different types of language units.To improve the performance of move recognition in scientific abstracts,a novel model of move recognition is proposed that outperforms the BERT-based method.Design/methodology/approach:Prevalent models based on BERT for sentence classification often classify sentences without considering the context of the sentences.In this paper,inspired by the BERT masked language model(MLM),we propose a novel model called the masked sentence model that integrates the content and contextual information of the sentences in move recognition.Experiments are conducted on the benchmark dataset PubMed 20K RCT in three steps.Then,we compare our model with HSLN-RNN,BERT-based and SciBERT using the same dataset.Findings:Compared with the BERT-based and SciBERT models,the F1 score of our model outperforms them by 4.96%and 4.34%,respectively,which shows the feasibility and effectiveness of the novel model and the result of our model comes closest to the state-of-theart results of HSLN-RNN at present.Research limitations:The sequential features of move labels are not considered,which might be one of the reasons why HSLN-RNN has better performance.Our model is restricted to dealing with biomedical English literature because we use a dataset from PubMed,which is a typical biomedical database,to fine-tune our model.Practical implications:The proposed model is better and simpler in identifying move structures in scientific abstracts and is worthy of text classification experiments for capturing contextual features of sentences.Originality/value:T he study proposes a masked sentence model based on BERT that considers the contextual features of the sentences in abstracts in a new way.The performance of this classification model is significantly improved by rebuilding the input layer without changing the structure of neural networks.Gaihong Yu Zhixiong Zhang Huan Liu Liangping Ding 2019Journal of Data and Information Science2019,4,4:14
3Big Data and Data Science:Opportunities and Challenges of iSchools显示文摘Due to the recent explosion of big data, our society has been rapidly going through digital transformation and entering a new world with numerous eye-opening developments. These new trends impact the society and future jobs, and thus student careers. At the heart of this digital transformation is data science, the discipline that makes sense of big data. With many rapidly emerging digital challenges ahead of us, this article discusses perspectives on iSchools' opportunities and suggestions in data science education. We argue that iSchools should empower their students with 'information computing' disciplines, which we define as the ability to solve problems and create values, information, and knowledge using tools in application domains. As specific approaches to enforcing information computing disciplines in data science education, we suggest the three foci of user-based, tool-based, and applicationbased. These three foci will serve to differentiate the data science education of iSchools from that of computer science or business schools. We present a layered Data Science Education Framework (DSEF) with building blocks that include the three pillars of data science (people, technology, and data), computational thinking, data-driven paradigms, and data science lifecycles. Data science courses built on the top of this framework should thus be executed with user-based, tool-based, and application-based approaches. This framework will help our students think about data science problems from the big picture perspective and foster appropriate problem-solving skills in conjunction with broad perspectives of data science lifecycles. We hope the DSEF discussed in this article will help fellow iSchools in their design of new data science curricula.Il-Yeol Song Yongjun Zhu 2017Journal of Data and Information Science2017,2,3:12
4Patent Citations Analysis and Its Value in Research Evaluation: A Review and a New Approach to Map Technology-relevant Research显示文摘Purpose:First,to review the state-of-the-art in patent citation analysis,particularly characteristics of patent citations to scientific literature(scientific non-patent references,SNPRs).Second,to present a novel mapping approach to identify technology-relevant research based on the papers cited by and referring to the SNPRs.Design/methodology/approach:In the review part we discuss the context of SNPRs such as the time lags between scientific achievements and inventions.Also patent-to-patent citation is addressed particularly because this type of patent citation analysis is a major element in the assessment of the economic value of patents.We also review the research on the role of universities and researchers in technological development,with important issues such as universities as sources of technological knowledge and inventor-author relations.We conclude the review part of this paper with an overview of recent research on mapping and network analysis of the science and technology interface and of technological progress in interaction with science.In the second part we apply new techniques for the direct visualization of the cited and citing relations of SNPRs,the mapping of the landscape around SNPRs by bibliographic coupling and co-citation analysis,and the mapping of the conceptual environment of SNPRs by keyword co-occurrence analysis.Findings:We discuss several properties of SNPRs.Only a small minority of publications covered by the Web of Science or Scopus are cited by patents,about 3%–4%.However,for publications based on university-industry collaboration the number of SNPRs is considerably higher,around 15%.The proposed mapping methodology based on a 'second order SNPR approach' enables a better assessment of the technological relevance of research.Research limitations:The main limitation is that a more advanced merging of patent and publication data,in particular unification of author and inventor names,in still a necessity.Practical implications:The proposed mapping methodology enables the creation of a database of technology-relevant papers(TRPs).In a bibliometric assessment the publications of research groups,research programs or institutes can be matched with the TRPs and thus the extent to which the work of groups,programs or institutes are relevant for technological development can be measured.Originality/value:The review part examines a wide range of findings in the research of patent citation analysis.The mapping approach to identify a broad range of technologyrelevant papers is novel and offers new opportunities in research evaluation practices.Anthony F.J. van Raan 2017Journal of Data and Information Science2017,2,1:12
5A Bibliometric Framework for Identifying'Princes'Who Wake up the'Sleeping Beauty'in Challenge-type Scientific Discoveries显示文摘Purpose:This paper develops and validates a bibliometric framework for identifying the 'princes'(PR) who wake up the 'sleeping beauty'(SB) in challenge-type scientific discoveries,so as to figure out the awakening mechanisms,and promote potentially valuable but not readily accepted innovative research.(A PR is a research study.)Design/methodology/approach:We propose that PR candidates must meet the following four criteria:(1) be published near the time when the SB began to attract a lot of citations;(2) be highly cited papers themselves;(3) receive a substantial number of co-citations with the SB;and(4) within the challenge-type discoveries which contradict established theories,the 'pulling effect' of the PR on the SB must be strong.We test the usefulness of the bibliometric framework through a case study of a key publication by the 2014 chemistry Nobel laureate Stefan W.Hell,who negated Ernst Abbe's diffraction limit theory,one of the most prominent paradigms in the natural sciences.Findings:The first-ranked candidate PR article identified by the bibliometric framework is in line with historical facts.An SB may need one or more PRs and even 'retinues' to be 'awakened.' Documents with potential awakening functionality tend to be published in prestigious multidisciplinary journals with higher impact and wider scope than the journals publishing SBs.Research limitations:The above framework is only applicable to transformative innovations,and the conclusions are drawn from the analysis of one typical SB and her awakening process.Therefore the generality of our work might be limited.Practical implications:Publications belonging to so-called transformative research,even when less frequently cited,should be given special attention as early as possible,because they may suddenly attract many citations after a period of sleep,as reflected in our case study.Originality/value:The definition of PR(s) as the first paper(s) that cited the SB article(selfciting excluded) has its limitations.Instead,the SB-PR co-citations should be given priority in current environment of scholarly communication.Since the 'premature' or 'transformative' breakthroughs in the challenge-type SB documents are either beyond the current knowledge domain,or violate established paradigms,people's psychological distance from the SB is larger than that from the PR,which explains why the annual citations of the PR are usually higher than those of the SB,especially prior to or during the SB's citation boom period.Jian Du YishanWu 2016Journal of Data and Information Science2016,1,1:9
6Smart Data for Digital Humanities显示文摘The emergence of 'Big Data' has been a dramatic development in recent years.Alongside it,a lesser-known but equally important set of concepts and practices has also come into being—'Smart Data.' This paper shares the author's understanding of what,why,how,who,where,and which data in relation to Smart Data and digital humanities.It concludes that,challenges and opportunities co-exist,but it is certain that Smart Data,the ability to achieve big insights from trusted,contextualized,relevant,cognitive,predictive,and consumable data at any scale,will continue to have extraordinary value in digital humanities.Marcia Lei Zeng 2017Journal of Data and Information Science2017,2,1:8
7Visualization of Disciplinary Profiles: Enhanced Science Overlay Maps显示文摘Purpose: The purpose of this study is to modernize previous work on science overlay maps by updating the underlying citation matrix, generating new clusters of scientific disciplines, enhancing visualizations, and providing more accessible means for analysts to generate their own maps.Design/methodology/approach: We use the combined set of 2015 Journal Citation Reports for the Science Citation Index (n of journals = 8,778) and the Social Sciences Citation Index (n=3,212) for a total of 11,365 journals. The set of Web of Science Categories in the Science Citation Index and the Social Sciences Citation Index increased from 224 in 2010 to 227 in 2015. Using dedicated software, a matrix of 227 × 227 cells is generated on the basis of whole-number citation counting. We normalize this matrix using the cosine function. We first develop the citing-side, cosine-normalized map using 2015 data and VOSviewer visualization with default parameter values. A routine for making overlays on the basis of the map('wc15.exe') is available at http://gffzzcc9cb02d7b4f4bb1h95f5pubfxqp966xn.ffgz.tsg.suse.edu.cn/wc15/index.htm.Findings: Findings appear in the form of visuals throughout the manuscript. In Figures 1–9 we provide basemaps of science and science overlay maps for a number of companies, universities, and technologies.Research limitations: As Web of Science Categories change and/or are updated so is the need to update the routine we provide. Also, to apply the routine we provide users need access to the Web of Science.Practical implications: Visualization of science overlay maps is now more accurate and true to the 2015 Journal Citation Reports than was the case with the previous version of the routine advanced in our paper.Originality/value: The routine we advance allows users to visualize science overlay maps in VOSviewer using data from more recent Journal Citation Reports.Stephen Carley Alan L.Porter Ismael Rafols Loet Leydesdorff 2017Journal of Data and Information Science2017,2,3:7
8A Criteria-based Assessment of the Coverage of Scopus and Web of Science显示文摘Purpose: The purpose of this study is to assess the coverage of the scientific literature in Scopus and Web of Science from the perspective of research evaluation.Design/methodology/approach: The academic communities of Norway have agreed on certain criteria for what should be included as original research publications in research evaluation and funding contexts. These criteria have been applied since 2004 in a comprehensive bibliographic database called the Norwegian Science Index(NSI). The relative coverages of Scopus and Web of Science are compared with regard to publication type, field of research and language.Findings: Our results show that Scopus covers 72 percent of the total Norwegian scientific and scholarly publication output in 2015 and 2016, while the corresponding figure for Web of Science Core Collection is 69 percent. The coverages are most comprehensive in medicine and health(89 and 87 percent) and in the natural sciences and technology(85 and 84 percent). The social sciences(48 percent in Scopus and 40 percent in Web of Science Core Collection) and particularly the humanities(27 and 23 percent) are much less covered in the two international data sources. Research limitation: Comparing with data from only one country is a limitation of the study, but the criteria used to define a country's scientific output as well as the identification of patterns of field-dependent partial representations in Scopus and Web of Science should be recognizable and useful also for other countries. Originality/value: The novelty of this study is the criteria-based approach to studying coverage problems in the two data sources.Dag W.Aksnes Gunnar Sivertsen 2019Journal of Data and Information Science2019,4,1:7
9Topic Detection Based on Weak Tie Analysis: A Case Study of LIS Research显示文摘Purpose: Based on the weak tie theory, this paper proposes a series of connection indicators Acof weak tie subnets and weak tie nodes to detect research topics, recognize their connections, and understand their evolution.Design/methodology/approach: First, keywords are extracted from article titles and preprocessed. Second, high-frequency keywords are selected to generate weak tie co-occurrence networks. By removing the internal lines of clustered sub-topic networks, we focus on the analysis of weak tie subnets' composition and functions and the weak tie nodes' roles.Findings: The research topics' clusters and themes changed yearly; the subnets clustered with technique-related and methodology-related topics have been the core, important subnets for years; while close subnets are highly independent, research topics are generally concentrated and most topics are application-related; the roles and functions of nodes and weak ties are diversified.Research limitations: The parameter values are somewhat inconsistent; the weak tie subnets and nodes are classified based on empirical observations, and the conclusions are not verified or compared to other methods.Practical implications: The research is valuable for detecting important research topics as well as their roles, interrelations, and evolution trends. Originality/value: To contribute to the strength of weak tie theory, the research translates weak and strong ties concepts to co-occurrence strength, and analyzes weak ties' functions. Also, the research proposes a quantitative method to classify and measure the topics' clusters and nodes.Ling Wei Haiyun Xu Zhenmeng Wang Kun Dong Chao Wang Shu Fang 2016Journal of Data and Information Science2016,1,4:6
10Rediscovering Don Swanson:The Past,Present and Future of Literature-based Discovery显示文摘Purpose: The late Don R.Swanson was well appreciated during his lifetime as Dean of the Graduate Library School at University of Chicago,as winner of the American Society for Information Science Award of Merit for 2000,and as author of many seminal articles.In this informal essay,I will give my personal perspective on Don's contributions to science,and outline some current and future directions in literature-based discovery that are rooted in concepts that he developed.Design/methodology/approach: Personal recollections and literature review.Findings: The Swanson A-B-C model of literature-based discovery has been successfully used by laboratory investigators analyzing their findings and hypotheses.It continues to be a fertile area of research in a wide range of application areas including text mining,drugrepurposing,studies of scientific innovation,knowledge discovery in databases,and bioinformatics.Recently,additional modes of discovery that do not follow the A-B-C model have also been proposed and explored(e.g.so-called storytelling,gaps,analogies,link prediction,negative consensus,outliers,and revival of neglected or discarded research questions).Research limitations: This paper reflects the opinions of the author and is not a comprehensive nor technically based review of literature-based discovery.Practical implications: The general scientific public is still not aware of the availability of tools for literature-based discovery.Our Arrowsmith project site maintains a suite of discovery tools that are free and open to the public(http://gffzz31a0be1597de4bbah95f5pubfxqp966xn.ffgz.tsg.suse.edu.cn),as does BITOLA which is maintained by Dmitar Hristovski(http://http://gffzz58b759aedd624224h95f5pubfxqp966xn.ffgz.tsg.suse.edu.cn/bitola),and Epiphanet which is maintained by Trevor Cohen(http://gffzz61b0bff3167f41c7h95f5pubfxqp966xn.ffgz.tsg.suse.edu.cn/).Bringing user-friendly tools to the public should be a high priority,since even more than advancing basic research in informatics,it is vital that we ensure that scientists actually use discovery tools and that these are actually able to help them make experimental discoveries in the lab and in the clinic.Originality/value: This paper discusses problems and issues which were inherent in Don's thoughts during his life,including those which have not yet been fully taken up and studied systematically.Neil R.Smalheiser 2017Journal of Data and Information Science2017,2,4:6
11CiteOpinion: Evidence-based Evaluation Tool for Academic Contributions of Research Papers Based on Citing Sentences显示文摘Purpose:To uncover the evaluation information on the academic contribution of research papers cited by peers based on the content cited by citing papers,and to provide an evidencebased tool for evaluating the academic value of cited papers.Design/methodology/approach:CiteOpinion uses a deep learning model to automatically extract citing sentences from representative citing papers;it starts with an analysis on the citing sentences,then it identifies major academic contribution points of the cited paper,positive/negative evaluations from citing authors and the changes in the subjects of subsequent citing authors by means of Recognizing Categories of Moves(problems,methods,conclusions,etc.),and sentiment analysis and topic clustering.Findings:Citing sentences in a citing paper contain substantial evidences useful for academic evaluation.They can also be used to objectively and authentically reveal the nature and degree of contribution of the cited paper reflected by citation,beyond simple citation statistics.Practical implications:The evidence-based evaluation tool CiteOpinion can provide an objective and in-depth academic value evaluation basis for the representative papers of scientific researchers,research teams,and institutions.Originality/value:No other similar practical tool is found in papers retrieved.Research limitations:There are difficulties in acquiring full text of citing papers.There is a need to refine the calculation based on the sentiment scores of citing sentences.Currently,the tool is only used for academic contribution evaluation,while its value in policy studies,technical application,and promotion of science is not yet tested.Xiaoqiu Le Jingdan Chu Siyi Deng Qihang Jiao Jingjing Pei Liya Zhu Junliang Yao 2019Journal of Data and Information Science2019,4,4:6
12Big Metadata,Smart Metadata,and Metadata Capital:Toward Greater Synergy Between Data Science and Metadata显示文摘Purpose:The purpose of the paper is to provide a framework for addressing the disconnect between metadata and data science. Data science cannot progress without metadata research.This paper takes steps toward advancing the synergy between metadata and data science, and identifies pathways for developing a more cohesive metadata research agenda in data science.Design/methodology/approach: This paper identifies factors that challenge metadata research in the digital ecosystem, defines metadata and data science, and presents the concepts big metadata, smart metadata, and metadata capital as part of a metadata lingua franca connecting to data science. Findings: The 'utilitarian nature' and 'historical and traditional views' of metadata are identified as two intersecting factors that have inhibited metadata research. Big metadata, smart metadata, and metadata capital are presented as part of a metadata lingua franca to help frame research in the data science research space.Research limitations: There are additional, intersecting factors to consider that likely inhibit metadata research, and other significant metadata concepts to explore.Practical implications: The immediate contribution of this work is that it may elicit response, critique, revision, or, more significantly, motivate research. The work presented can encourage more researchers to consider the significance of metadata as a research worthy topic within data science and the larger digital ecosystem. Originality/value: Although metadata research has not kept pace with other data science topics, there is little attention directed to this problem. This is surprising, given that metadata is essential for data science endeavors. This examination synthesizes original and prior scholarship to provide new grounding for metadata research in data science.Jane Greenberg 2017Journal of Data and Information Science2017,2,3:6
13Data-driven Discovery: A New Era of Exploiting the Literature and Data显示文摘In the current data-intensive era, the traditional hands-on method of conducting scientific research by exploring related publications to generate a testable hypothesis is well on its way of becoming obsolete within just a year or two. Analyzing the literature and data to automatically generate a hypothesis might become the de facto approach to inform the core research efforts of those trying to master the exponentially rapid expansion of publications and datasets. Here, viewpoints are provided and discussed to help the understanding of challenges of data-driven discovery.Ying Ding Kyle Stirling 2016Journal of Data and Information Science2016,1,4:6
14Functions of Uni- and Multi-citations: Implications for Weighted Citation Analysis显示文摘Purpose:(1) To test basic assumptions underlying frequency-weighted citation analysis:(a) Uni-citations correspond to citations that are nonessential to the citing papers;(b) The influence of a cited paper on the citing paper increases with the frequency with which it is cited in the citing paper.(2) To explore the degree to which citation location may be used to help identify nonessential citations.Design/methodology/approach:Each of the in-text citations in all research articles published in Issue 1 of the Journal of the Association for Information Science and Technology(JASIST) 2016 was manually classified into one of these five categories:Applied,Contrastive,Supportive,Reviewed,and Perfunctory.The distributions of citations at different in-text frequencies and in different locations in the text by these functions were analyzed.Findings:Filtering out nonessential citations before assigning weight is important for frequency-weighted citation analysis.For this purpose,removing citations by location is more effective than re-citation analysis that simply removes uni-citations.Removing all citation occurrences in the Background and Literature Review sections and uni-citations in the Introduction section appears to provide a good balance between filtration and error rates.Research limitations:This case study suffers from the limitation of scalability and generalizability.We took careful measures to reduce the impact of other limitations of the data collection approach used.Relying on the researcher's judgment to attribute citation functions,this approach is unobtrusive but speculative,and can suffer from a low degree of confidence,thus creating reliability concerns.Practical implications:Weighted citation analysis promises to improve citation analysis for research evaluation,knowledge network analysis,knowledge representation,and information retrieval.The present study showed the importance of filtering out nonessential citations before assigning weight in a weighted citation analysis,which may be a significant step forward to realizing these promises.Originality/value:Weighted citation analysis has long been proposed as a theoretical solution to the problem of citation analysis that treats all citations equally,and has attracted increasing research interest in recent years.The present study showed,for the first time,the importance of filtering out nonessential citations in weighted citation analysis,pointing research in this area in a new direction.Dangzhi Zhao Alicia Cappello Lucinda Johnston 2017Journal of Data and Information Science2017,2,1:5
15Methods and Practices for Institutional Benchmarking based on Research Impact and Competitiveness: A Case Study of ShanghaiTech University显示文摘Purpose: To develop and test a mission-oriented and multi-dimensional benchmarking method for a small scale university aiming for internationally first-class basic research.Design/methodology/approach: An individualized evidence-based assessment scheme was employed to benchmark ShanghaiTech University against selected top research institutions,focusing on research impact and competitiveness at the institutional and disciplinary levels.Topic maps opposing ShanghaiTech and corresponding top institutions were produced for the main research disciplines of ShanghaiTech. This provides opportunities for further exploration of strengths and weakness. Findings: This study establishes a preliminary framework for assessing the mission of the university. It further provides assessment principles, assessment questions, and indicators.Analytical methods and data sources were tested and proved to be applicable and efficient.Research limitations: To better fit the selective research focuses of this university, its schema of research disciplines needs to be re-organized and benchmarking targets should include disciplinary top institutions and not necessarily those universities leading overall rankings.Current reliance on research articles and certain databases may neglect important research output types.Practical implications: This study provides a working framework and practical methods for mission-oriented, individual, and multi-dimensional benchmarking that ShanghaiTech decided to use for periodical assessments. It also offers a working reference for other institutions to adapt. Further needs are identified so that ShanghaiTech can tackle them for future benchmarking.Originality/value: This is an effort to develop a mission-oriented, individually designed,systematically structured, and multi-dimensional assessment methodology which differs from often used composite indices.Jiang Chang Jianhua Liu 2019Journal of Data and Information Science2019,4,3:5
16A Study of Methods to Identify Industry-University-Research Institution Cooperation Partners based on Innovation Chain Theory显示文摘Purpose: This study aims at identifying potential industry-university-research collaboration(IURC) partners effectively and analyzes the conditions and dynamics in the IURC process based on innovation chain theory.Design/methodology/approach: The method utilizes multisource data, combining bibliometric and econometrics analyses to capture the core network of the existing collaboration networks and institution competitiveness in the innovation chain. Furthermore, a new identification method is constructed that takes into account the law of scientific research cooperation and economic factors.Findings: Empirical analysis of the genetic engineering vaccine field shows that through the distribution characteristics of creative technologies from different institutions, the analysis based on the innovation chain can identify the more complementary capacities among organizations.Research limitations: In this study, the overall approach is shaped by the theoretical concept of an innovation chain, a linear innovation model with specific types or stages of innovation activities in each phase of the chain, and may, thus, overlook important feedback mechanisms in the innovation process.Practical implications: Industry-university-research institution collaborations are extremely important in promoting the dissemination of innovative knowledge, enhancing the quality of innovation products, and facilitating the transformation of scientific achievements.Originality/value: Compared to previous studies, this study emulates the real conditions of IURC. Thus, the rule of technological innovation can be better revealed, the potential partners of IURC can be identified more readily, and the conclusion has more value.Haiyun Xu Chao Wang Kun Dong Rui Luo Zenghui Yue Hongshen Pang 2018Journal of Data and Information Science2018,3,2:5
17The Flemish Performance-based Research Funding System: A Unique Variant of the Norwegian Model显示文摘The BOF-key is the performance-based research funding system that is used in Flanders, Belgium. In this paper we describe the historical background of the system, its current design and organization, as well as its effects on the Flemish higher education landscape. The BOFkey in its current form relies on three bibliometric parameters: publications in Web of Science, citations in Web of Science, and publications in a comprehensive regional database for SSH publications. Taken together, the BOF-key forms a unique variant of the Norwegian model: while the system to a large extent relies on a commercial database, it avoids the problem of inadequate coverage of the SSH. Because the bibliometric parameters of the BOF-key are reused in other funding allocation schemes, their overall importance to the Flemish universities is substantial.Tim C.E.Engels Raf Guns 2018Journal of Data and Information Science2018,3,4:4
18Topic Evolution and Emerging Topic Analysis Based on Open Source Software显示文摘Purpose:We present an analytical,open source and flexible natural language processing and text mining method for topic evolution,emerging topic detection and research trend forecasting for all kinds of data-tagged text.Design/methodology/approach:We make full use of the functions provided by the open source VOSviewer and Microsoft Office,including a thesaurus for data clean-up and a LOOKUP function for comparative analysis.Findings:Through application and verification in the domain of perovskite solar cells research,this method proves to be effective.Research limitations:A certain amount of manual data processing and a specific research domain background are required for better,more illustrative analysis results.Adequate time for analysis is also necessary.Practical implications:We try to set up an easy,useful,and flexible interdisciplinary text analyzing procedure for researchers,especially those without solid computer programming skills or who cannot easily access complex software.This procedure can also serve as a wonderful example for teaching information literacy.Originality/value:This text analysis approach has not been reported before.Xiang Shen Li Wang 2020Journal of Data and Information Science2020,5,4:4
19Embedding-based Detection and Extraction of Research Topics from Academic Documents Using Deep Clustering显示文摘Purpose:Detection of research fields or topics and understanding the dynamics help the scientific community in their decisions regarding the establishment of scientific fields.This also helps in having a better collaboration with governments and businesses.This study aims to investigate the development of research fields over time,translating it into a topic detection problem.Design/methodology/approach:To achieve the objectives,we propose a modified deep clustering method to detect research trends from the abstracts and titles of academic documents.Document embedding approaches are utilized to transform documents into vector-based representations.The proposed method is evaluated by comparing it with a combination of different embedding and clustering approaches and the classical topic modeling algorithms(i.e.LDA)against a benchmark dataset.A case study is also conducted exploring the evolution of Artificial Intelligence(AI)detecting the research topics or sub-fields in related AI publications.Findings:Evaluating the performance of the proposed method using clustering performance indicators reflects that our proposed method outperforms similar approaches against the benchmark dataset.Using the proposed method,we also show how the topics have evolved in the period of the recent 30 years,taking advantage of a keyword extraction method for cluster tagging and labeling,demonstrating the context of the topics.Research limitations:We noticed that it is not possible to generalize one solution for all downstream tasks.Hence,it is required to fine-tune or optimize the solutions for each task and even datasets.In addition,interpretation of cluster labels can be subjective and vary based on the readers’opinions.It is also very difficult to evaluate the labeling techniques,rendering the explanation of the clusters further limited.Practical implications:As demonstrated in the case study,we show that in a real-world example,how the proposed method would enable the researchers and reviewers of the academic research to detect,summarize,analyze,and visualize research topics from decades of academic documents.This helps the scientific community and all related organizations in fast and effective analysis of the fields,by establishing and explaining the topics.Originality/value:In this study,we introduce a modified and tuned deep embedding clustering coupled with Doc2Vec representations for topic extraction.We also use a concept extraction method as a labeling approach in this study.The effectiveness of the method has been evaluated in a case study of AI publications,where we analyze the AI topics during the past three decades.Sahand Vahidnia Alireza Abbasi Hussein A.Abbass 2021Journal of Data and Information Science2021,6,3:4
20Identifying Scientific Project-generated Data Citation from Full-text Articles: An Investigation of TCGA Data Citation显示文摘Purpose: In the open science era, it is typical to share project-generated scientific data by depositing it in an open and accessible database. Moreover, scientific publications are preserved in a digital library archive. It is challenging to identify the data usage that is mentioned in literature and associate it with its source. Here, we investigated the data usage of a government-funded cancer genomics project, The Cancer Genome Atlas(TCGA), via a full-text literature analysis.Design/methodology/approach: We focused on identifying articles using the TCGA dataset and constructing linkages between the articles and the specific TCGA dataset. First, we collected 5,372 TCGA-related articles from Pub Med Central(PMC). Second, we constructed a benchmark set with 25 full-text articles that truly used the TCGA data in their studies, and we summarized the key features of the benchmark set. Third, the key features were applied to the remaining PMC full-text articles that were collected from PMC.Findings: The amount of publications that use TCGA data has increased significantly since 2011, although the TCGA project was launched in 2005. Additionally, we found that the critical areas of focus in the studies that use the TCGA data were glioblastoma multiforme, lung cancer, and breast cancer; meanwhile, data from the RNA-sequencing(RNA-seq) platform is the most preferable for use.Research limitations: The current workflow to identify articles that truly used TCGA data is labor-intensive. An automatic method is expected to improve the performance.Practical implications: This study will help cancer genomics researchers determine the latest advancements in cancer molecular therapy, and it will promote data sharing and data-intensive scientific discovery.Originality/value: Few studies have been conducted to investigate data usage by governmentfunded projects/programs since their launch. In this preliminary study, we extracted articles that use TCGA data from PMC, and we created a link between the full-text articles and the source data.Jiao Li Si Zheng Hongyu Kang Zhen Hou Qing Qian 2016Journal of Data and Information Science2016,1,2:4
返回顶部 每页显示:
共19页 首页 上一页 第1页 下一页 末页 /19 跳转

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费