|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | Big Earth data:A new frontier in Earth and information sciences显示文摘Big data is a revolutionary innovation that has allowed the development of many new methods in scientific research.This new way of thinking has encouraged the pursuit of new discoveries.Big data occupies the strategic high ground in the era of knowledge economies and also constitutes a new national and global strategic resource.“Big Earth data”,derived from,but not limited to,Earth observation has macro-level capabilities that enable rapid and accurate monitoring of the Earth,and is becoming a new frontier contributing to the advancement of Earth science and significant scientific discoveries.Within the context of the development of big data,this paper analyzes the characteristics of scientific big data and recognizes its great potential for development,particularly with regard to the role that big Earth data can play in promoting the development of Earth science.On this basis,the paper outlines the Big Earth Data Science Engineering Project(CASEarth)of the Chinese Academy of Sciences Strategic Priority Research Program.Big data is at the forefront of the integration of geoscience,information science,and space science and technology,and it is expected that big Earth data will provide new prospects for the development of Earth science. | Huadong Guo | 2017 | Big Earth Data2017,1,1: | 49 |
| 2 | Multi-view Clustering: A Survey显示文摘In the big data era, the data are generated from different sources or observed from different views. These data are referred to as multi-view data. Unleashing the power of knowledge in multi-view data is very important in big data mining and analysis. This calls for advanced techniques that consider the diversity of different views,while fusing these data. Multi-view Clustering(MvC) has attracted increasing attention in recent years by aiming to exploit complementary and consensus information across multiple views. This paper summarizes a large number of multi-view clustering algorithms, provides a taxonomy according to the mechanisms and principles involved, and classifies these algorithms into five categories, namely, co-training style algorithms, multi-kernel learning, multiview graph clustering, multi-view subspace clustering, and multi-task multi-view clustering. Therein, multi-view graph clustering is further categorized as graph-based, network-based, and spectral-based methods. Multi-view subspace clustering is further divided into subspace learning-based, and non-negative matrix factorization-based methods. This paper does not only introduce the mechanisms for each category of methods, but also gives a few examples for how these techniques are used. In addition, it lists some publically available multi-view datasets.Overall, this paper serves as an introductory text and survey for multi-view clustering. | Yan Yang Hao Wang | 2018 | Big Data Mining and Analytics2018,1,2: | 20 |
| 3 | Mapping landslide susceptibility and types using Random Forest显示文摘Landslides are one of the most destructive natural hazards;they can drastically alter landscape morphology,destroy man-made struc-tures,and endanger people’s life.Landslide susceptibility maps(LSMs),which show the spatial likelihood of landslide occurrence,are crucial for environmental management,urban planning,and minimizing economic losses.To date,the majority of research into data mining LSM uses small-scale case studies focusing on a single type of landslide.This paper presents a data mining approach to producing LSM for a large,heterogeneous region that is susceptible tomultipletypesoflandslides.UsingacasestudyofPiedmont,Italy,a Random Forest algorithm is applied to produce both susceptibility maps and classification maps.These maps are combined to give a highly accurate(over 85%classification accuracy)LSM which con-tains a large amount of information and is easy to interpret.This novel method of mapping landslide susceptibility demonstrates the efficacy of Random Forest to produce highly accurate susceptibility maps for alargeheterogeneousregion withouttheneed formultiple susceptibility assessments. | Khaled Taalab Tao Cheng Yang Zhang | 2018 | Big Earth Data2018,2,2: | 13 |
| 4 | Big data drives the development of Earth science显示文摘Big data is now a popular topic,becoming increasingly well known around the world,and yet the concept of big data and its implications are still novel.To discuss big data,it is appropriate to first talk about what is really meant by this term,and so to begin the first article in the inaugural issue of Big Earth Data,let us look at how data has become big data and why that is important. | Huadong Guo | 2017 | Big Earth Data2017,1,1: | 11 |
| 5 | DEEPEYE: Link Prediction in Dynamic Networks Based on Non-negative Matrix Factorization显示文摘A Non-negative Matrix Factorization(NMF)-based method is proposed to solve the link prediction problem in dynamic graphs. The method learns latent features from the temporal and topological structure of a dynamic network and can obtain higher prediction results. We present novel iterative rules to construct matrix factors that carry important network features and prove the convergence and correctness of these algorithms. Finally, we demonstrate how latent NMF features can express network dynamics efficiently rather than by static representation,thereby yielding better performance. The amalgamation of time and structural information makes the method achieve prediction results that are more accurate. Empirical results on real-world networks show that the proposed algorithm can achieve higher accuracy prediction results in dynamic networks in comparison to other algorithms. | Nahla Mohamed Ahmed Ling Chen Yulong Wang Bin Li Yun Li Wei Liu | 2018 | Big Data Mining and Analytics2018,1,1: | 11 |
| 6 | Relation Classification via Recurrent Neural Network with Attention and Tensor Layers显示文摘Relation classification is a crucial component in many Natural Language Processing(NLP) systems. In this paper, we propose a novel bidirectional recurrent neural network architecture(using Long Short-Term Memory,LSTM, cells) for relation classification, with an attention layer for organizing the context information on the word level and a tensor layer for detecting complex connections between two entities. The above two feature extraction operations are based on the LSTM networks and use their outputs. Our model allows end-to-end learning from the raw sentences in the dataset, without trimming or reconstructing them. Experiments on the SemEval-2010 Task 8dataset show that our model outperforms most state-of-the-art methods. | Runyan Zhang Fanrong Meng Yong Zhou Bing Liu | 2018 | Big Data Mining and Analytics2018,1,3: | 9 |
| 7 | Generation of ready to use (RTU) products over China based on Landsat series data显示文摘Earth observation community has entered into the era of big data.Family of Landsat sensors have collected massive medium resolution satellite images,which are valuable for long-term land surface monitoring.In order to significantly reduce the magnitude of data processing for remote sensing data users,Landsat-based Ready to Use(RTU)products have been produced.Main RTU products,including orthorectified products,land surface reflectance,land surface temperature,large-area mosaic image,and standard image map products,are described.The resulting Landsat RTU products are hosted on the RSGS earth observation data sharing web site for free download(http://gffzz959b1a8e74bf485dhvvkp0uo6opxq6b0o.ffgz.tsg.suse.edu.cn/rtu/).These new products will provide consistent,standardized,multi-decadal image data for robust land cover change detection and monitoring across the Earth sciences.In the coming years,CASEarth DataBank system will be constructed,which is an intelligent data service platform for providing not only the RTU products from multi-source satellite data,but also big earth data analysis methods. | Guojin He Zhaoming Zhang Weili Jiao Tengfei Long Yan Peng Guizhou Wang Ranyu Yin Wei Wang Xiaomei Zhang Huichan Liu Bo Cheng Bo Xiang | 2018 | Big Earth Data2018,2,1: | 8 |
| 8 | A Novel Deep Hybrid Recommender System Based on Auto-encoder with Neural Collaborative Filtering显示文摘Due to the widespread availability of implicit feedback(e.g., clicks and purchases), some researchers have endeavored to design recommender systems based on implicit feedback. However, unlike explicit feedback,implicit feedback cannot directly reflect user preferences. Therefore, although more challenging, it is also more practical to use implicit feedback for recommender systems. Traditional collaborative filtering methods such as matrix factorization, which regards user preferences as a linear combination of user and item latent vectors, have limited learning capacities and suffer from data sparsity and the cold-start problem. To tackle these problems,some authors have considered the integration of a deep neural network to learn user and item features with traditional collaborative filtering. However, there is as yet no research combining collaborative filtering and contentbased recommendation with deep learning. In this paper, we propose a novel deep hybrid recommender system framework based on auto-encoders(DHA-RS) by integrating user and item side information to construct a hybrid recommender system and enhance performance. DHA-RS combines stacked denoising auto-encoders with neural collaborative filtering, which corresponds to the process of learning user and item features from auxiliary information to predict user preferences. Experiments performed on the real-world dataset reveal that DHA-RS performs better than state-of-the-art methods. | Yu Liu Shuai Wang M.Shahrukh Khan Jieyu He | 2018 | Big Data Mining and Analytics2018,1,3: | 7 |
| 9 | Big Earth data facilitates sustainable development goals显示文摘The 2030 Agenda for Sustainable Development,comprising 17 Sustainable Development Goals(SDGs),was adopted in September 2015 by heads of state and government at the United Nations(UN)summit.The Agenda is a transformative plan of action for people,planet and prosperity that all countries and all stakeholders will implement to ensure that no one is left behind.The inception of the 2030 Agenda marks a milestone in the progress towards a sustainable society for all. | Huadong Guo | 2020 | Big Earth Data2020,4,1: | 7 |
| 10 | The challenges of a Big Data Earth显示文摘The potential of big data fused with the vision of a digital Earth offers powerful opportunities to deepen understanding of the whole Earth system and the management of a sustainable planet.It is important to stand back from often confusing detail to clarify what those opportunities are and how they might be seized.The essential scientific potential of data,big or small,is to reveal patterns,which have often been the fundamental first step in stimulating inquiry,leading to new questions,new perspectives and potentially to new answers.The digital revolution has created a“digital microscope”that permits us to see patterns that have not been seen before,and when coupled with machine learning technologies to analyse them in creating statistical predictions of the behaviour of both human and non-human systems.These potentials converge with the imperative to represent an Earth system with interacting non-human and human components,as a vital contribution to the understanding and actions required in working towards planetary sustainability.But a digital Earth is also capable of being represented mathematically as a digitally networked phenomenon,analogous to an analogue computer,and should be an important target for a Big Earth Data Journal.We should also return to Al Gore’s vision of an accessible digital Earth with wide usability.Pre-determining the separate functions of parallel digital Earths risks losing one of the great potentials of big data and learning algorithms,the identification and analysis of unanticipated relationships and processes. | Geoffrey Boulton | 2018 | Big Earth Data2018,2,1: | 7 |
| 11 | MODIS-based Daily Lake Ice Extent and Coverage dataset for Tibetan Plateau显示文摘The Tibetan Plateau houses numerous lakes,the phenology and duration of lake ice in this region are sensitive to regional and global climate change,and as such are used as key indicators in climate change research,particularly in environment change comparison studies for the Earth three poles.However,due to its harsh natural environment and sparse population,there is a lack of conventional in situ measurement on lake ice phenology.The Moderate Resolution Imaging Spectroradiometer(MODIS)Normalized Difference Snow Index(NDSI)data,which can be traced back 20 years with a 500 m spatial resolution,were used to monitor lake ice for filling the observation gaps.Daily lake ice extent and coverage under clear-sky conditions was examined by employing the conventional SNOWMAP algorithm,and those under cloud cover conditions were re-determined using the temporal and spatial continuity of lake surface conditions through a series of steps.Through time series analysis of every single lake with size greater than 3 km2 in size,308 lakes within the Tibetan Plateau were identified as the effective records of lake ice extent and coverage to form the Daily Lake Ice Extent and Coverage dataset,including 216 lakes that can be further retrieved with four determinable lake ice parameters:Freeze-up Start(FUS),Freeze-up End(FUE),Break-up Start(BUS),and Break-up End(BUE),and 92 lakes with two parameters,FUS and BUE.Six lakes of different sizes and locations were selected for verification against the published datasets by passive microwave remote sensing.The lake ice phenology information obtained in this paper was highly consistent with that from passive microwave data at an average correlation coefficient of 0.91 and an RMSE value varying from 0.07 to 0.13.The present dataset is more effective at detecting lake ice parameters for smaller lakes than the coarse resolution passive microwave remote sensing observations.The published data are available in http://gffzz324c53c67c144612svvkp0uo6opxq6b0o.ffgz.tsg.suse.edu.cn/repository/uuid:fdfd8c76-6b7c-4bbf-aec8-98ab199d9093 and http://gffzzc14d9fd2be724b39hvvkp0uo6opxq6b0o.ffgz.tsg.suse.edu.cn/dataSet/handle/744. | Yubao Qiu Pengfei Xie Matti Leppäranta Xingxing Wang Juha Lemmetyinen Hui Lin Lijuan Shi | 2019 | Big Earth Data2019,3,2: | 7 |
| 12 | Building an Earth Observations Data Cube: lessons learned from the Swiss Data Cube (SDC) on generating Analysis Ready Data (ARD)显示文摘Pressures on natural resources are increasing and a number of challenges need to be overcome to meet the needs of a growing population in a period of environmental variability.Some of these environmental issues can be monitored using remotely sensed Earth Observations(EO)data that are increasingly available from a number of freely and openly accessible repositories.However,the full information potential of EO data has not been yet realized.They remain still underutilized mainly because of their complexity,increasing volume,and the lack of efficient processing capabilities.EO Data Cubes(DC)are a new paradigm aiming to realize the full potential of EO data by lowering the barriers caused by these Big data challenges and providing access to large spatio-temporal data in an analysis ready form.Systematic and regular provision of Analysis Ready Data(ARD)will significantly reduce the burden on EO data users.Nevertheless,ARD are not commonly produced by data providers and therefore getting uniform and consistent ARD remains a challenging task.This paper presents an approach to enable rapid data access and pre-processing to generate ARD using interoperable services chains.The approach has been tested and validated generating Landsat ARD while building the Swiss Data Cube. | Gregory Giuliani Bruno Chatenoux Andrea De Bono Denisa Rodila Jean-Philippe Richard Karin Allenbach Hy Dao Pascal Peduzzi | 2017 | Big Earth Data2017,1,1: | 7 |
| 13 | Atmospheric heat source/sink dataset over the Tibetan Plateau based on satellite and routine meteorological observations显示文摘The Tibetan Plateau(TP),acting as a large elevated land surface and atmospheric heat source during spring and summer,has a substantial impact on regional and global weather and climate.To explore the multi-scale temporal variation in the thermal forcing effect of the TP,here we calculated the surface sensible heat and latent heat release based on 6-h routine observations at 80(32)meteorological stations during the period 1979–2016(1960–2016).Meanwhile,in situ air-column net radiation cooling during the period 1984–2015 was derived from satellite data.This new data-set provides continuous,robust,and the longest observational atmospheric heat source/sink data over the third pole,which will be helpful to better understand the spatial-temporal structure and multi-scale variation in TP diabatic heating and its influence on the earth’s climatic system. | Anmin Duan Senfeng Liu Yu Zhao Kailun Gao Wenting Hu | 2018 | Big Earth Data2018,2,2: | 7 |
| 14 | Innovative approaches to the Sustainable Development Goals using Big Earth Data显示文摘A persistent challenge for the Sustainable Development Goals(SDGs)has been a lack of data for indicators to assess progress towards each goal and varying capacities among nations to con-duct these assessments.Rapid developments in big data,however,are facilitating a global approach to the SDGs.Tools and data products are emerging that can be extended to and leveraged by nations that do not yet have the capacity to measure SDG indica-tors.Big Earth Data,a special class of big data,integrates multisource data within a geographic context,utilizing the principles and methodologies of the established literature on big data science,applied specifically to Earth system science.This paper discusses the research challenges related to Big Earth Data and the concerted efforts and investments required to make and mea-sure progress towards the SDGs.As an example,the Big Earth Data Science Engineering Program(CASEarth)of the Chinese Academy of Sciences is presented along with other case studies on Big Earth Data in support of the SDGs.Lastly,the paper proposes future priorities for developments in Big Earth Data,such as human resource capacity,digital infrastructure,interoperability,and envir-onmental considerations. | Huadong Guo Dong Liang Fang Chen Zeeshan Shirazi | 2021 | Big Earth Data2021,5,3: | 7 |
| 15 | Big Data Analytics for Healthcare Industry:Impact,Applications,and Tools显示文摘In recent years, huge amounts of structured, unstructured, and semi-structured data have been generated by various institutions around the world and, collectively, this heterogeneous data is referred to as big data. The health industry sector has been confronted by the need to manage the big data being produced by various sources,which are well known for producing high volumes of heterogeneous data. Various big-data analytics tools and techniques have been developed for handling these massive amounts of data, in the healthcare sector. In this paper, we discuss the impact of big data in healthcare, and various tools available in the Hadoop ecosystem for handling it. We also explore the conceptual architecture of big data analytics for healthcare which involves the data gathering history of different branches, the genome database, electronic health records, text/imagery, and clinical decisions support system. | Sunil Kumar Maninder Singh | 2019 | Big Data Mining and Analytics2019,2,1: | 6 |
| 16 | Location Prediction on Trajectory Data: A Review显示文摘Location prediction is the key technique in many location based services including route navigation, dining location recommendations, and traffic planning and control, to mention a few. This survey provides a comprehensive overview of location prediction, including basic definitions and concepts, algorithms, and applications. First, we introduce the types of trajectory data and related basic concepts. Then, we review existing location-prediction methods, ranging from temporal-pattern-based prediction to spatiotemporal-pattern-based prediction. We also discuss and analyze the advantages and disadvantages of these algorithms and briefly summarize current applications of location prediction in diverse fields. Finally, we identify the potential challenges and future research directions in location prediction. | Ruizhi Wu Guangchun Luo Junming Shao Ling Tian Chengzong Peng | 2018 | Big Data Mining and Analytics2018,1,2: | 5 |
| 17 | A Novel Clustering Technique for Efficient Clustering of Big Data in Hadoop Ecosystem显示文摘Big data analytics and data mining are techniques used to analyze data and to extract hidden information.Traditional approaches to analysis and extraction do not work well for big data because this data is complex and of very high volume. A major data mining technique known as data clustering groups the data into clusters and makes it easy to extract information from these clusters. However, existing clustering algorithms, such as k-means and hierarchical, are not efficient as the quality of the clusters they produce is compromised. Therefore, there is a need to design an efficient and highly scalable clustering algorithm. In this paper, we put forward a new clustering algorithm called hybrid clustering in order to overcome the disadvantages of existing clustering algorithms. We compare the new hybrid algorithm with existing algorithms on the bases of precision, recall, F-measure, execution time, and accuracy of results. From the experimental results, it is clear that the proposed hybrid clustering algorithm is more accurate, and has better precision, recall, and F-measure values. | Sunil Kumar Maninder Singh | 2019 | Big Data Mining and Analytics2019,2,4: | 5 |
| 18 | Exploring the depths of the global earth observation system of systems显示文摘This paper explores for the first time the contents,structure and relationships across institutions and disciplines of a global Big Earth Data cyber-infrastructure:the Global Earth Observation System of System(GEOSS).The analysis builds on 1.8 million metadata records harvested in GEOSS.Because this set includes almost all the major large data collections in GEOSS,the analysis represents more than 80%of all the data made available through this global system.We explore two major aspects:the collaborative networks and the thematic coverage in GEOSS.The first connects the contributing organisations through the more than 200,000 keywords used in the systems,and then explores who is citing whom,a proxy for of institutional thickness.The thematic coverage is analysed through neural network algorithms,first on the keywords,and then on the corpus of 653 million lemmatised lower case words built from the titles and abstracts of all 1.8 million metadata records.The findings not only give a good overview of the GEOSS data universe,but offer immediate priorities on how to increase the usability of GEOSS through improved data management,and the opportunity to augment the metadata with high level concept that synthetise well the contents of the data-set. | Max Craglia Jiri Hradec Stefano Nativi Mattia Santoro | 2017 | Big Earth Data2017,1,1: | 5 |
| 19 | A global land cover map produced through integrating multi-source datasets显示文摘In the past decades,global land cover datasets have been produced but also been criticized for their low accuracies,which have been affecting the applications of these datasets.Producing a new global dataset requires a tremendous amount of efforts;however,it is also possible to improve the accuracy of global land cover mapping by fusing the existing datasets.A decision-fuse method was developed based on fuzzy logic to quantify the consistencies and uncertainties of the existing datasets and then aggregated to provide the most certain estimation.The method was applied to produce a 1-km global land cover map(SYNLCover)by integrating five global land cover datasets and three global datasets of tree cover and croplands.Efforts were carried out to assess the quality:1)inter-comparison of the datasets revealed that the SYNLCover dataset had higher consistency than these input global land cover datasets,suggesting that the data fusion method reduced the disagreement among the input datasets;2)quality assessment using the human-interpreted reference dataset reported the highest accuracy in the fused SYNLCover dataset,which had an overall accuracy of 71.1%,in contrast to the overall accuracy between 48.6%and 68.9%for the other global land cover datasets. | Min Feng Yan Bai | 2019 | Big Earth Data2019,3,3: | 5 |
| 20 | An Improved Hybrid Collaborative Filtering Algorithm Based on Tags and Time Factor显示文摘The Collaborative Filtering(CF) recommendation algorithm, one of the most popular algorithms in Recommendation Systems(RS), mainly includes memory-based and model-based methods. When performing rating prediction using a memory-based method, the approach used to measure the similarity between users or items can significantly influence the recommendation performance. Traditional CFs suffer from data sparsity when making recommendations based on a rating matrix, and cannot effectively capture changes in user interest. In this paper, we propose an improved hybrid collaborative filtering algorithm based on tags and a time factor(TTHybridCF), which fully utilizes tag information that characterizes users and items. This algorithm utilizes both tag and rating information to calculate the similarity between users or items. In addition, we introduce a time weighting factor to measure user interest, which changes over time. Our experimental results show that our method alleviates the sparsity problem and demonstrates promising prediction accuracy. | Chunxia Zhang Ming Yang Jing Lv Wanqi Yang | 2018 | Big Data Mining and Analytics2018,1,2: | 4 |