Weblog

From messy data to linked data: LOD-enabled Google Refine

Some of you are probably already familiar with Google Refinehttp://code.google.com/p/google-refine/ (GR), a simple yet very powerful tool for working with messy data. GR is a web application, based on a modular web application framework, which makes it an interesting blend of versatility, performance, portability, simplicity and extendability. GR has no database, all data is in the in-memory data-store, which optimized for the operations used by faceted browsing – one of the most useful features when dealing with messy data; another one is support for reconciliation services. However, currently GR’s data cleansing, reconciliation abilities and user interface are mostly limited to working with subsets of FreebaseFreebase is a large collaborative knowledge base. It is an online collection of structured data harvested from many sources, including individual 'wiki' contribution. Freebase aims to create a global resource which allows people (and machines) to access common information more effectively. It is ... entities. Wouldn’t it be nice to be able to use DBpediaDBpedia is a project aiming to extract structured information from the information created as part of the Wikipedia project. This structured information is then made available on the World Wide Web. DBpedia allows users to query relationships and properties associated with Wikipedia resources, ..., too?

Now it is possible: we made GoogleGoogle Inc. is a multinational public corporation invested in Internet search, cloud computing, and advertising technologies. Google hosts and develops a number of Internet-based services and products, and generates profit primarily from advertising through its AdWords program. The company was ... Refine LOD-friendly with GR extensions provided by DERICentre for Science, Engineering and Technology (CSET) established in 2003 with funding from the Science Foundation Ireland. The vision of the Digital Enterprise Research Institute is to be recognised as one of the leading international web science research institutes interlinking technologies, ... and Zemanta. DERI’s RDF extension takes care of reconciliation with any SPARQL point or a RDF dump, and it does the trick when data needs to be exported into RDF. On the other hand, Zemanta’s DBpedia extension adds two more LOD-related functionalities: the first one is augmentation of reconciled data with additional data from DBpedia and the second one is extraction of entities in full text by entity type to new columns using ZemantaZemanta is a content suggestion engine for bloggers and other content creators. APIAn application programming interface (API) is an interface implemented by a software program to enable interaction with other software, similar to the way a user interface facilitates interaction between humans and computers. APIs are implemented by applications, libraries and operating systems ....

LOD-enabled Google Refine (LODGrefine) is also available as a package with pre-integrated extensions mentioned above. The source code of extensions and the package is available on Github under the BSD Licence: RDF extension, DBpedia extension and LODGrefine.

Stay tuned – we have more plans with LODGrefine. First, LODGrefine will be integrated into LOD2EU-funded (FP7) research project aiming to take the Web of Linked Data to the next level. Main research challenges: improve coherence and quality of data published on the Web, close the performance gap between relational and RDF data management, establish trust on the Linked Data Web and ... Stack, then we’ll explore the possibility of integrating crowd-sourcing solution like Amazon Mechanical Turk… but more about this when the time comes.

Good news for the end: LODGrefine will be presented at SemTech2012 conference in San Francisco in June. Hope to see you there.

Enhanced by Zemanta

Posted in Announcement, Misc, WP4 – Reuse, Interlinking and Knowledge Fusion

One Response to From messy data to linked data: LOD-enabled Google Refine

  1. Pingback: Catch me if you can: uncaught TypeError undefined, Twitter bootstrap and LODGrefine « Sparkica's Brainmachine

Leave a Reply

Your email address will not be published. Required fields are marked *

*

You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <strike> <strong>


join our monthly webinars

Public Mailinglist & Newsletter

Please subscribe me to the LOD2 mailinglist.
my email address
my name (optional)
goto archive

Follow


Follow lod2project on Twitter

RSS Twitter

  • linkeddata
    RT @igorop: #mashpoint http://t.co/eTPcq2NV now in browser mode allows iterative query refinement, #databrowser #semweb #linkeddata […]
  • linkeddata
    RT @BigDataClub: RT @nopiedra: “@BigDataPVSW: Solving the #bigdata problem with #RDF http://t.co/5iv65pBn #linkeddata […]
  • linkeddata
    RT @BigDataClub: RT @nopiedra: “@BigDataPVSW: Solving the #bigdata problem with #RDF http://t.co/5iv65pBn #linkeddata […]