Implementation of search feature should cover three main areas, which are indexing, searching and presentation. Such separation gives us modularity, which implies reuse of components and ability to test easily.
Indexing would be backed by Apache Lucene indexing mechanisms.
Indexing should be performed in two phases:
Phase one: gathering basic data
Every file in contribution should be parsed and indexed. Moreover files contained in every known archive file (JAR, ZIP files) should be indexed. Parsing each file would be aware of its filetype to get best indexing data. Internally, Document object would be created and appropriate fields would be filled. Unique identifier should be assigned to newly created document.
Phase two: reference completion
For every indexed document previously created index will be searched to find documents usages. Found elements would be stored with current document data. In example, for indexed Java class we should search for its name occurences in composite files. Found identifiers as well as their friendly names would be stored in indexed document representing Java class.
Index document model
Each indexed item would have several attributes. Bold items would be document fields for searching purposes. Italic is used for attributes used for internal purposes.
Field |
Description |
|---|---|
Filename |
Full path for contribution object. |
Friendly name |
specific for file type, like QName of Java class, Qname of composite. For others it could be simply filename. |
File content |
literal for text files, for non readable files, ie. Java classes some names like method names could be extracted and used in this field |
Contribution |
contribution which file belongs to |
Archive |
archive which file belongs to |
References |
Contains names of items which are used by current: |
Usages |
Contains names of items which uses current. Generally it's reversed link for reference. |
Service/Reference/ |
Used to store extracted objects from composite file. |
All |
All above fields to provide non-filter queries. |
Item identifier |
Unique across SCA domain to identify document/domain obejct. |
References links |
Links to references documents. |
Usages links |
Links to usages documents. |
Above model shows how documents would be linked for search purposes. Example of references which could occur could be found on the diagram.
Searching would be backed by Apache Lucene search engine.
Custom search API for Apache Tuscany would be available via SCA component. Such element can be reused in various scenarios, ie. it can be exposed for other purposes via one of Apache Tuscany bindings. In this project we would like to use such component as a feed for web based UI.
There would be generally two operations exposed by such component. (more could be introduced for ie. administration purposes).
Fetch by phrase - search phrases are similiar to what we do in Google. Lucene query syntax would be used and various filters could be entered (basing on fields described in 2.2.1 Indexing). User could filter by:
More query syntax elements would be used, such as:
Fetch by item - getting item and its references. Item to fetch is identified by internal identifier. Such fetching method would be used in navigation based on hyperlinks, not search queries.
Navigation
Navigation could be performed in two ways:
1. By using search box where user can type query, for "fetch by phrase" search method
2. By using links to items where user can navigate through references ("fetch by item" method). Such links could be found in several places:
Additionally after implementing project core some usability features coulb be added:
Results
Display layout would be common for both "fetch by phrase" and "fetch by item". Every search would be displayed as list of results. For long result lists paging would be applied. Furthermore having sort (basing on various criteria) feature would improve navigation through results list.
View for each search result element should contain:
Following image shows example navigation throught search UI. It contains 5 web pages which can be reach in various flows. Red arrows shows what page wuold be generated after clicking a link. Purple color is used for comments.

Contribution scanner, parser and analyzer
Module which scans contributions, analyzes its artifacts and feeds Apache Lucene index. See 2.2.1 Indexing for details. Appriopriate JUnit tests should be introduced.
Search component
Module which exposes search features as simple, SCA component API. See 2.2.2 Searching for details. Appropriate JUnit tests should be introduced.
SCA Domain manager web application extension
New pages, scripts, actions etc. which would handle UI described in 2.2.3 Presentation. Appropriate JUnit testst should be introduced.
Integration tests
Module which tests integration of project deliverables with Apache Tuscany.
Documentation
User and developer documentation on Apache Tuscany web page.
Getting started
Proposal review and discussions, prototyping, getting familiar with advanced aspects of related technologies.
Implementation of Contribution scanner, parser and analyzer.
Implementation of Search component.
Implementation of Integration tests.
Submitting mid-term evaluation.
Implementation of SCA Domain manager web application extension.
Extra week in case of delay.
Writing documentation for project. Code review.
Submitting final evaluation