Semantic Web Company

The Semantic Puzzle

Open World Assumptions

subscribe RSS

Kingsley Idehen: “By declaring its context, Linked Data can be made more easily reusable by others”

June 16, 2010 By: Andreas Blumauer Category: Corporate Semantic Web, Enterprise 2.0, Linked Data & Open Data, Tools & Software No Comments →

Semantic Web Company talked with Kingsley Idehen who is CEO of OpenLink Software and probably one of the most profound experts on data integration issues about “Linked Data”.

The interview covers questions like:

  • How can Linked Data help to make companies more productive?
  • Do you think that the Linked Data Initiative can build upon a stable architecture or will it face more and more problems the bigger the “cloud” will grow?
  • What´s the ultimate argument for an Enterprise Architect to use languages like SPARQL at least in addition to SQL?
  • How will a “Real Time Semantic Web” change the whole game?
  • How will the “Semantic Web” be called in 10 years? Will there still be a “Semantic Web”?

Read the full version of the interview here.

Sphere: Related Content

Linked Data is not owl:sameAs Semantic Web

March 30, 2009 By: Andreas Blumauer Category: Linked Data & Open Data, Search Engines 3 Comments →

twitter_cloudletWhile some people work heavily on the extension of the semantic web infrastructure, like Talis Connected Commons or OpenLink´s Amazon EC2 Instantiation others have started to bring the semantic web closer to the developers and therefore to a much broader audience: They offer search facilities or Linked Data Navigators like OpenLink´s Entity Finder or DERI´s VisiNav.

Those kind of applications should not be confused with “semantic web” end-user-applications like Google´s Wonderwheel or INTSPEI´s Cloudlet: To add some semantics to existing user-interfaces can be helpful and obviously users are ready for such experiments, but of course this is NOT the innovation which the semantic web will bring but it is a very important step to be taken in parallel with the linked data initiative.

Let´s take a look at Cloudlet: This tool is an easy-to-use free Firefox extension that adds context-sensitive tag clouds to the most popular search engines and helps people more efficiently navigate through their search results. The previous version of Search Cloudlet worked with Google and Yahoo; the new version also works with Twitter. It adds Tag Clouds, Author Clouds, Recipient Clouds and Hashtag Clouds to Twitter search, Twitter user profiles and home pages. See some reviews on this popular tool.

Cloudlet is a child of the Web. INTSPEI has learned all lessons from Web 2.0 especially how to promote ideas using the blogosphere and how to identify market trends as early as possible, and it generates some added value for the users which is obvious. Sure, it doesn´t make use of linked data yet, but as a typical representative of the fast growing “semantic search evolution” it reminds me on Chris Welty´s famous insight: “In the Semantic Web, it is not the Semantic which is new, it is the Web which is new.”

Web 1.0 was the WWW without tons of network effects. Web 2.0 changed that a lot.

Linked Data is not the Semantic Web, it´s the basement for it. From a software developer´s and an IT archictect´s perspective it might seem as those two concepts were the same. But this community represents a very small percentage of all web-users.

So where is the User´s Web in the Linked Data architecture? If you´re looking at TimBL´s Linked Data principles one can clearly see that this is a “Web” for developers.

But things evolve. And some Web companies will jump on the bandwagon and will, for instance, improve their tagclouds, their semantic search, their recommender systems (Twine?) or their similarity search a lot by making use of linked data.

Like semantic search becomes mainstream (or call it “semantic search 2.0″) right now, then (in about three years, I guess) linked data will become part of a lot of mainstream applications. Linked data will generate tons of new network effects, maybe even new business models, it won´t be avant-garde anymore. It will be part of the Semantic Web.

Sphere: Related Content

DBpedia, UMBEL & the Future Web’s Ecology – interview with Mike Bergman & Sören Auer

November 10, 2008 By: Andreas Blumauer Category: Linked Data & Open Data, Mashups & Web services, Ontology Engineering 5 Comments →

Sören AuerThe Linked Open Data infrastructure is in a tremendous process of maturing – the recent release of UMBEL’s webservice AND the incorporation of UMBEL classes in DBpedia are yet another confirmation of this exciting process. Knowing and having met DBpedia co-initiator, Triplify main developer and head of the AKSW research group Sören Auer and UMBEL editor and Zitgist CEO Mike Bergman in various contexts, I felt it was time to talk to and pick the brains of both these key players in a dialog situation. The (first) result is the interview you can find below. As not everyone can expected to be familiar with both projects, here is some backgrond to get you started (you can also go directly to the interview):

Sören Auer (image above), Mike Bergman (image below)

DBpedia has become the largest RDF repository for encyclopaedic knowledge, extracting structured information from Wikipedia and making it available on the Web of Data. UMBEL, on the other hand, provides an OpenCYC-based, light-weight ontology structure for relating Web content and data to a standard set of subject concepts, with a number of 20,000 concepts currently reached. In the Linked Data Cloud, DBpedia and UMBEL map and cross-reference each other.

Mike BergmanIn practice this means that UMBEL provides classes to describe the concepts to which “things” are members. For instance, named entities from Wikipedia such as “John F. Kennedy” are mapped with subject concepts such as Leader, Person, Administrator and Graduate, with broader and equivalent classes in CYC and FOAF and broader subject concepts within UMBEL. A link is set to Wikipedia, as well as a ‘same as’ reference to DBpedia. A class structure enables faceted browsing and extraction, inferencing, and navigation and discovery for all datasets linked to that structure.

DBpedia, in turn, returns properties of ‘John J. Kennedy’ (e.g. abstracts in available Wikipedia languages, demographic information such as birth date and place, alma mater, predecessors and successors), and ‘same as’ references, e.g., to the JFK entry in Freebase (who recently released their RDF service) and the aforementioned page in UMBEL. Furthermore, DBpedia maps the URI with available RDF types, for instance foaf:person or yago:AssassinatedAmericanPoliticians and, once again, with UMBEL’s subject concepts Person, Administrator, Graduate and Leader.

Due to its reliance on Wikipedia, DBpedia does a great job at covering a bandwidth of knowledge as broad as the spectrum of the interest of people participating in Wikipedia; it’s within the area of named entities, i.e. entities such as persons, organizations, locations, which have a proper name, but are not necessarily and specifically part of a particular, acknowledged domain or discipline. UMBEL, on the other hand, has as its most apparent advantage its reliance on OpenCyc and with that the strong inferencing and logic capabilities of the CYC knowledge-base which are thus also brought to the Web of Data. DBpedia is a community project started by the University of Leipzig, Free University Berlin and OpenLink Software, while the open and free UMBEL is developed and hosted by Zitgist with support from, again, OpenLink Software.

Now, and in particular with the recent release of Zitgist’s web service endpoints and with the incorporation of UMBEL classes in DBpedia, questions arises as to the relationship of the two projects, and regarding the role of OpenLink Software in the further process. To draw a distinction:

One could say that DBpedia’s goal is to lower the barrier for web developers and end-users in the actual use of the semantic web, while UMBEL aims at bringing “order to the chaos” that is inherent to user-generated, collective knowledge.

Would you agree with this description – and is it a contradiction at all or the kind of dynamic the Semantic Web community has been waiting for?

Mike Bergman: Yes, I would agree with this description, though we have tried many others. For example, in various writings in the past, we have described UMBEL as a roadmap, or middleware, or a backbone, or a concept ontology, or an ‘infocline’, or a meta layer for metadata, and others. Today, what I tend to use, particularly in reference to DBpedia, is the TBox-ABox distinction in computer science and description logics. UMBEL is more of a class or structural and concept relationships schema — a TBox — while DBpedia is more of an an instance and entity layer with attributes — an ABox. I think they are pretty complementary…
(more…)

Sphere: Related Content

Bringing (Legacy) Data to the Web [WOD-PD]

October 22, 2008 By: Jana Herwig Category: Conferences & Events, Linked Data & Open Data 2 Comments →

The third session at WOD-PD was dedicated to “Bringing (Legacy) Data on the Web“, and led by Sören Auer (University of Leipzig, Germany) and Orri Erling (OpenLink Software) .

Sören Auer giving a talkSören Auer described the difference between the Web 1.0, 2.0 and 3.0 as follows: On the Web 1.0, you had many websites that provided unstructured, mainly textual content. On the Web 2.0, you have a few large websites that are specialised on specific content types. And, finally, on the Web 3.0, there are many websites which contain, and are able to semantically syndicate, arbitrarily structured content.

So why would we need another web? What you cannot do with the current web is finding answers to seemingly complex, yet in reality pretty mundane question such as: Where in Leipzig do I find an apartment that is close to bilingual, German-French child care facilities? Are there any ERP service providers which have offices in Vienna and Berlin? Who are the researchers in South-East Asia currently working on database related topics?

Sören further discussed three of the present means of bringing relation data to the web: Triplify (a web application plugin that exposes data from relational databases in RDF), D2RQ (a declarative language to describe mappings between relational database schemata and OWL/RDFS ontologies, developed at Free University Berlin), and Virtuoso Universal Server (a middleware and database engine hybrid delivering for instance data integration for SQL, RDF, XML, Web Services). With respect to Triplify, Sören – who is Triplify’s founder and main developer at AKSW Uni Leipzig – showed and discussed the configuration for Wordpress 2.1., which can be found here (click here for more configurations, e.g. for Joomla, OpenConf and Drupal). The next aim for Triplify is to become an integral part in enduser web app distibutions.

And important question raised by Sören was: How do next generation search engines know that something has changed on the web of data? He suggested three approaches:

  1. Always try to crawl everything (this may sound silly – but that’s actually what is happening on the current web)
  2. Ping a central update notification service – e.g. PingTheSemanticWeb.com – which works as a showcase, but will probably not scale if the data web gets really deployed.
  3. Each linked data endpoint publishes an update log – e.g. with Triplify, as a special folder inside the Triplify namespace, e.g. http://example.com/Triplify/update

Also discussed by Sören and worth checking out is Reuters’ Semantic proxy – the demo went live in late September.

Orri Erling, as the lead developer of the Virtuoso Team, addressed the issue of mapping relational databases to RDF with OpenLink Virtuoso. In his talk, he addressed the pros and cons of RDF data warehouse:

Pros

  • Even query performance across all data
  • Possibility of forward-chaining inference
  • Some SPARQL features may be better supported, e.g. Unspecified predicates

Cons

  • Keeping data up-to-date
  • Complex set up, needs dedicated servers: you don’t build them on a whim

Orri Erling giving a talkWhat Virtuoso delivers is mapping of SPARQL to SQL against any existing schema (whether stored in Virtuoso or elsewhere); a physical quad-store (quad as in quadruple; not as in quad-bike :) ; and Federated/local Relational Data Base Management Systems (RDBMS).

A more detailed discussion of the requirements for Relational-to-RDF Mapping is available on Orri’s blog, where he discusses it in the light of his own experience. A power point presentation of a previous talk he gave to the W3C RDB2RDF Incubator Group can be downloaded here: Mapping Relational Databases to RDF with OpenLink Virtuoso (PPT, 115KB). His summary of the group discussions around the same topic, Requirements for Relational to RDF Mapping, can be found here.

Orri also showed the Virtuoso billion triples demo which, according to the corresponding blogpost, “is being worked on at the time of submission and may be shown online by appointment.” The demo was a submission to the Billion Triples Challenge.

Reblog this post [with Zemanta]
Sphere: Related Content

LinkedData Planet in New York: A great community event for all things semantic

June 18, 2008 By: Andreas Blumauer Category: Conferences & Events 1 Comment →

Roosevelt HotelFirst of all: LinkedData Planet is a big success in terms of visitor numbers to begin with. The “Grand Ballroom” at Hotel Roosevelt, a lovely old hotel in Manhattan, was packed not only when Tim Berners-Lee gave his keynote this afternoon.

But not only the quantity of attendees, also the high quality of talks and discussions which are going on at this first conference on the “commercialization of the web of Linked Data” show that we are facing a fast growing phenomenon with a “great momentum” as Berners-Lee stated.

Kingsley Idehen from OpenLink Software started the first day of the conference with his keynote in which he tried to “demystify” the term Linked Data. He said that “Linked Data is the foundation of the semantic web, its connectivity is growing and the line between enterprise and individual level is blurring”. He also stressed the similarities between ODBC (Open DataBase Connectivity) and Linked Data – which might be interesting for my next talk with an “old-fashioned” CTO.

Uche Ogbuji from Zepheira referred to DBpedia as “the star” of the Linked Data Cloud and gave an interesting talk about the possibilities of Linking Enterprise Data.

Tim Berners-Lee listed in his keynote the areas in which the LOD-community is now facing the biggest challenges:

  • standards, for instance levels of inference, link following on linked data clients and servers
  • federated query; query service descriptions
  • ultimate human interface to all the data there are
  • balancing diversity & harmony in ontology development
  • and, of course, continuing the great momentum

Tim Berners-Lee also emphasised that hiding information in some cases like product information is “just crazy”. One way to expand the LOD cloud could be “lobbying” for data sources (with governments, providers of commercial information, etc).

Between the talks people spent their time in the exhibition area. Dean Allemang from TopQuadrant gave a demo of TopBraid Composer. I talked to Tom Tague from OpenCalais about the new features the next releases will have and how to use their service behind the firewall and Mike Bergman from Zitgist showed me the power of UMBEL web services.

All in all – LinkedData Planet is a great community event, well organised and well populated by people who want to use the semantic web in different commercial settings.

And to those who weren’t able to attend: I recommend to take a look at the Linked Data Shopping List, a page within the Linked Data Initiative’s wiki where you can add the data that you want to see published as Linked Data.

Read also pt. 2 of our conference report: The social hub @ LinkedData Planet 2008

[Image: official-ly cool]

Zemanta Pixie
Sphere: Related Content

LinkedData Planet – Conference & Expo 2008

April 17, 2008 By: Jana Herwig Category: Conferences & Events 3 Comments →

Come share your expertise with linked data and semantic technologies and learn from others at LinkedData Planet in New York City (June 17-18, 2008).

In creating the modern generation of enterprise and web applications, we typically integrate information from multiple sources. Relating data from disparate sources presents a challenge of deriving information. However, semantic tools and technologies are evolving that enable us to understand information derived by linking data from different sources, including data from applications, databases, ontologies and content management systems. Semantic technologies and tools support techniques such as tagging online information to make it more readily accessible for data integration. This makes it easier to understand data in relation to other data, even if some of this data is inside your firewall, some is in a business partner’s system, and some is part of the growing collection of useful publicly available data on the web.

LinkedData Planet provides insights into those technologies that enableus to:

  • connect data contained in silos within organizations in a meaningful way
  • extract and correlate data from web sites and databases for purposes such as analyzing trends and decision support, customer and vendor relationship management, and social networking

(more…)

Sphere: Related Content