· 11 years ago · Sep 05, 2015, 09:21 PM
1Network Access to Multimedia Information
2
3 33
4
5 Network Access to Multimedia Information
6
7
8
9 Chris Adie
10
11 Edinburgh University Computing Service
12 University Library Building
13 George Square
14 Edinburgh
15 EH8 9LJ
16 Great Britain
17
18
19
20 Second Edition - 9 August, 1993
21
22
23
24 RARE Project OBR(93)015
25
26 Réseaux Associés pour la Recherche Européenne
27 Singel 466-468, NL-1017 AW
28 AMSTERDAM
29 Netherlands
30
31 Copyright RARE 1993
32
33 Abstract
34
35This report summarises the requirements of research and
36academic network users for network access to multimedia
37information. It does this by investigating some of the
38projects planned or currently underway in the community.
39Existing information systems such as Gopher, WAIS and
40World-Wide Web are examined from the point of view of
41multimedia support, and some interesting hypermedia
42systems emerging from the research community are also
43studied. Relevant existing and developing standards in
44this area are discussed. The report identifies the gaps
45between the capabilities of currently-deployed systems
46and the user requirements, and proposes further work
47centred on the World-Wide Web system to rectify this.
48
49The report is in some places very detailed, so it is
50preceded by an extended summary, which outlines the
51findings of the report.
52
53 Publication History
54
55The first edition was released on 29 June 1993. This
56second edition contains minor changes, corrections and
57updates. Change bars mark such revisions.
58
59 Extended Summary
60
61Introduction
62
63This report is concerned with issues in the intersection
64of networked information retrieval, database and
65multimedia technologies. It aims to establish research
66and academic user requirements for network access to
67multimedia data, to look at existing systems which offer
68partial solutions, and to identify what needs to be done
69to satisfy the most pressing requirements.
70
71User Requirements
72
73There are a number of reasons why multimedia data may
74need to be accessed remotely (as opposed to physically
75distributing the data, eg on CD-ROM). These reasons
76centre on the cost of physical distribution, versus the
77timeliness of network distribution. Of course, there is
78a cost associated with network distribution, but this
79tends to be hidden from the end user.
80
81User requirements have been determined by studying
82existing and proposed projects involving networked
83multimedia data. It has proved convenient to divide the
84applications into four classes according to their
85requirements: multimedia database applications, academic
86(particularly scientific) publishing applications, cal
87(computer-aided learning), and general multimedia
88information services.
89
90Database applications typically involve large collections
91of monomedia (non-text) data with associated textual and
92numeric fields. They require a range of search and
93retrieval techniques.
94
95Publishing applications require a range of media types,
96hyperlinking, and the capability to access the same data
97using different access paradigms (search, browse,
98hierarchical, links). Authentication and charging
99facilities are required.
100
101Cal applications require sophisticated presentation and
102synchronisation capabilities, of the type found in
103existing multimedia authoring tools. Authentication and
104monitoring facilities are required.
105
106General multimedia information services include on-line
107documentation, campus-wide information systems, and other
108systems which don't conveniently fall into the preceding
109categories. Hyperlinking is perhaps the most common
110requirement in this area.
111
112The analysis of these application areas allows a number
113of important user requirements to be identified:
114
115 Support for the Apple Macintosh, UNIX and PC/MS Windows
116 environments.
117
118 Support for a wide range of media types - text, image,
119 graphics and application-specific media being most
120 important, followed by video and sound.
121
122 Support for hyperlinking, and for multiple access
123 structures to be built on the same underlying data.
124
125 Support for sophisticated synchronisation and
126 presentation facilities.
127
128 Support for a range of database searching techniques.
129
130 Support for user annotation of information, and for
131 user-controlled display of sequenced media.
132
133 Adequate responsiveness - the maximum time taken to
134 retrieve a node should not exceed 20s.
135
136 Support for user authentication, a charging mechanism,
137 and monitoring facilities.
138
139 The ability to execute scripts.
140
141 Support for mail-based access to multimedia documents,
142 and (where appropriate) for printing multimedia
143 documents.
144
145 Powerful, easy-to-use authoring tools.
146
147Existing Systems
148
149The main information retrieval systems in use on the
150Internet are Gopher, Wais, and the World-Wide Web. All
151work on a client-server paradigm, and all provide some
152degree of support for multimedia data.
153
154Gopher presents the user with a hierarchical arrangement
155of nodes which are either directories (menus), leaf nodes
156(documents containing text or other media types), or
157search nodes (allowing some set of documents to be
158searched using keywords, possibly using WAIS). A range
159of media types is supported. Extensions currently being
160developed for Gopher (Gopher+) provide better support for
161multimedia data. Gopher has a very high penetration
162(there are over 1000 Gopher servers on the Internet), but
163it does not provide hyperlinks and is inflexibly
164hierarchical.
165
166Wais (Wide Area Information Server) allows users to
167search for documents in remote databases. Full-text
168indexing of the databases allows all documents containing
169particular (combinations of) words to be identified and
170retrieved. Non-text data (principally image data) can be
171handled, but indexing such documents is only performed on
172the document file name, severely limiting its usefulness.
173However, WAIS is ideally suited to text search
174applications.
175
176World-Wide Web (WWW) is a large-scale distributed
177hypermedia system. The Web consists of nodes (also
178called documents) and links. Links are connections
179between documents: to follow a link, the user clicks on a
180highlighted word in the source document, which causes the
181linked-to document to be retrieved and displayed. A
182document can be one of a variety of media types, or it
183can be a search node in a similar sense to Gopher. The
184WWW addressing method means that WAIS and Gopher servers
185may also be accessed from (indeed, form part of) the Web.
186WWW has a smaller penetration than Gopher, but is growing
187faster. The Web technology is currently being revised to
188take better account of the needs of multimedia
189information.
190
191These systems all go some way to meet the user
192requirements.
193
194 Support for multiple platforms and for a wide range of
195 media types (through "viewer" software external to the
196 client program) is good.
197
198 Only WWW has hyperlinks.
199
200 There is little or no support for sophisticated
201 presentation and synchronisation requirements.
202
203 Support for database querying tends to be limited to
204 "keyword" searches, but current developments in Gopher
205 and WWW should make more sophisticated queries
206 possible.
207
208 Some clients support user annotation of documents.
209
210 Response times for all three systems vary substantially
211 depending on the network distance between client and
212 server, and there is no support for isochronous data
213 transfer.
214
215 There is little in the way of authentication, charging
216 and monitoring facilities, although these are planned
217 for WWW.
218
219 Scripting is not supported because of security issues
220
221 WWW supports a mail responder.
222
223 The only system sufficiently complex to warrant an
224 authoring tool is WWW, which has editors to support its
225 hypertext markup language.
226
227Research
228
229There are a number of research projects which are of
230significant interest.
231
232Hyper-G is an ambitious distributed hypermedia research
233project at the University of Graz. It combines concepts
234of hypermedia, information retrieval systems and
235documentation systems with aspects of communication and
236collaboration, and computer-supported teaching and
237learning. Automatic generation of hyperlinks is
238supported, and there is a concept of generic structures
239which can exist in parallel with the hyperlink structure.
240Hyper-G is based on UNIX, and is in use as a CWIS at
241Graz. Gateways between Hyper-G and WWW exist.
242
243Microcosm is a PC-based hypermedia system developed at
244the University of Southampton. It can be viewed as an
245integrating hypermedia framework - a layer on top of a
246range of existing applications which enables
247relationships between different documents to be
248established. Hyperlinks are maintained separately from
249the data. Networking support for Microcosm is currently
250under development, as are versions of Microcosm for the
251Apple Macintosh and for UNIX. Microcosm is currently
252being "commercialised".
253
254AthenaMuse 2 is an ambitious distributed hypermedia
255authoring and presentation system under development by a
256university/industry consortium based at MIT. It will
257have good facilities for presentation and synchronisation
258of multimedia data, strong authoring support, and will
259include support for networking isochronous data. It will
260be a commercial product. Initial versions will support
261UNIX and X windows, with a PC/MS Windows version
262following. Apple Macintosh support has lower priority.
263
264The "Xanadu" project is designing and building an "open,
265social hypermedia" distributed environment, but shows no
266sign of delivering anything after several years of work.
267
268The European Commission sponsors a number of peripherally
269relevant projects through its Esprit and RACE research
270programmes. These programmes tend to be oriented towards
271commercial markets, and are thus not directly relevant.
272An exception is the Esprit IDOMENEUS project, which
273brings together workers in the database, information
274retrieval and multimedia fields. It is recommended that
275RARE establish a liaison with this project.
276
277There are a variety of other academic and commercial
278research projects which are also of interest. None of
279them are as directly relevant as those outlined above.
280
281Standards
282
283There are a number of existing and emerging standards for
284structuring hypermedia applications. Of these, the most
285important are SGML, HyTime, MHEG, ODA, PREMO and Acrobat.
286All bar the last are de jure standards, while Acrobat is
287a commercial product which is being proposed as a de
288facto standard.
289
290SGML (Standard Generalized Markup Language) is a markup
291language for delimiting the logical and semantic content
292of text documents. Because of its flexibility, it has
293become an important tool in hypermedia systems. HyTime
294is an ISO standardised infrastructure for representing
295integrated, open hypermedia documents, and is based on
296SGML. HyTime has great expressive power, but is not
297optimised for run-time efficiency. It is recommended
298that future RARE work on networked hypermedia should take
299account of the importance of SGML and HyTime.
300
301MHEG (Multimedia and Hypermedia information coding
302Experts Group) is a draft ISO standard for representing
303hypermedia applications in a platform-independent form.
304It uses an object-oriented approach, and is optimised for
305run-time efficiency. Full IS status for MHEG is expected
306in 1994. It is recommended that RARE keep a watching
307brief on MHEG.
308
309The ODA (Open Document Architecture) standard is being
310enhanced to incorporate multimedia and hypermedia
311features. However, interest in ODA is perceived to be
312decreasing, and it is recommended that ODA should not
313form a basis for further RARE work in networked
314hypermedia.
315
316PREMO is a new work item in the ISO graphics
317standardisation community, which appears to overlap with
318MHEG and HyTime. It is not clear that the PREMO work,
319which is at a very early stage, is worthwhile in view of
320the existence of those standards.
321
322Acrobat PDF is a format for representing multimedia
323(printable) documents in a portable, revisable form. It
324is based on Postscript, and is being proposed by Adobe
325Inc (originators of Postscript) as an industry standard.
326RARE should maintain awareness of this technology in view
327of its potential impact on multimedia information
328systems.
329
330There are various standards which have relevance to the
331way multimedia data is accessed across the network. Many
332of these have been described in a previous report
333[Adi93]. Two further access protocols are the proposed
334multimedia extensions to SQL, and the Document Filing and
335Retrieval protocol. Neither of these are likely to have
336major significance for networked multimedia information
337systems.
338
339Other standards of importance include:
340
341 MIME, a multimedia email standard which defines a range
342 of media types and encoding methods for those types
343 which are useful in a wider context.
344
345 AVIs (Audio-Visual Interactive services) and the
346 associated multimedia scripting language SMSL, which
347 form a standardisation initiative within CCITT (now ITU-
348 TSS) to specify interactive multimedia services which
349 can be provided across telephone/ISDN networks.
350
351There are two important trade associations which are
352involved in standardisation work. The Interactive
353Multimedia Association (IMA) has a Compatibility Project
354which is developing a specification for platform-
355independent interactive multimedia systems, including
356networking aspects. A newly-formed group, the Multimedia
357Communications Forum (MMCF), plans to provide input to
358the standards bodies. It is recommended that RARE become
359an Observing Member of the MMCF. A third trade
360association - the Multimedia Communications Community of
361Interest - has also just been formed.
362
363Future Directions
364
365Three common design approaches emerge from the variety of
366systems and standards analysed in this report. They can
367be described in terms of distinctions between different
368aspects of the system:
369
370 content is distinct from hyperstructure
371
372 media type is distinct from media encoding
373
374 data is distinct from protocol
375
376Distributed hypermedia systems are emerging from the
377research/development phase into the experimental
378deployment phase. However, the existing global
379information systems (Gopher, WAIS and WWW) are still
380largely limited to the use of external viewers for non-
381textual data. The most significant mismatches between
382the capabilities of currently-deployed systems and user
383requirements are in the areas of presentation and quality
384of service (ie responsiveness).
385
386Improving QOS is significantly more difficult than
387improving presentation capabilities, but there are a
388number of possible ways in which this could be addressed.
389Improving feedback to the user, greater multi-threading
390of applications, pre-fetching, caching, the use of
391alternative "views" of a node, and the use of isochronous
392data streams are all avenues which are worth exploring.
393
394In order to address these problems, it is recommended
395that RARE seek to adapt and enhance existing tools,
396rather than develop new ones. In particular, it is
397recommended that RARE select the World-Wide Web to
398concentrate its efforts on. The reasons for this choice
399revolve around the flexibility of the WWW design, the
400availability of hyperlinks, the existing effort which is
401already going into multimedia support in WWW, the fact
402that it is an integrating solution incorporating both
403WAIS and Gopher support, and its high rate of growth
404compared to Gopher (despite Gopher's wider deployment).
405Gopher is the main competitor to WWW, but its inflexibly
406hierarchical structure and the absence of hyperlinks make
407it difficult to use for highly-interactive multimedia
408applications.
409
410It is recommended that RARE should invite proposals for
411and subsequently commission work to:
412
413 Develop conversion tools from commercial multimedia
414 authoring packages to WWW, and accompanying authoring
415 guidelines.
416
417 Implement and evaluate the most promising ways of
418 overcoming the QOS problem.
419
420 Implement a specific user project using these tools, to
421 validate that the facilities being developed are truly
422 relevant to real applications.
423
424 Use the experience gained to inform and influence the
425 development of the WWW technology.
426
427 Contribute to the development of PC/MS Windows and
428 Apple Macintosh WWW clients, particularly in the
429 multimedia data handling area.
430
431It is noted that the rapid growth of WWW may in the
432future lead to problems through the implementation of
433multiple, uncoordinated and mutually incompatible add-on
434features. To guard against this trend, it may be
435appropriate for RARE, in coordination with CERN and other
436interested parties such as NCSA, to:
437
438 Encourage the formation of a consortium to coordinate
439 WWW technical development.
440
441 Table of Contents
442
443Acknowledgements 2
444Disclaimer 2
445Availability 2
4461. Introduction 3
447 1.1. Background 3
448 1.2. Terminology 3
4492. User Requirements 5
450 2.1. Applications 5
451 2.2. Data Characteristics 8
452 2.3. Requirements Definition 9
4533. Existing Systems 13
454 3.1. Gopher 13
455 3.2. Wide Area Information Server 17
456 3.3. World-Wide Web 19
457 3.4. Evaluating Existing Tools 24
4584. Research 29
459 4.1. Hyper-G 29
460 4.2. Microcosm 30
461 4.3. AthenaMuse 2 31
462 4.4. CEC Research Programmes 32
463 4.5. Other 33
4645. Standards 35
465 5.1. Structuring Standards 35
466 5.2. Access Mechanisms 40
467 5.3. Other Standards 40
468 5.4. Trade Associations 42
4696. Future Directions 44
470 6.1. General Comments on the State-of-the-Art
471 44
472 6.2. Quality of Service 44
473 6.3. Recommended Further Work 46
474References 49
475
476
477 Acknowledgements
478
479The following people have (knowingly or unknowingly)
480helped in the preparation of this report: Tim Berners-
481Lee, John Dyer, Aydin Edguer, Anton Eliens, Tony Gibbons,
482Stewart Granger, Wendy Hall, Gary Hill, Brian Marquardt,
483Gunnar Moan, Michael Neuman, Ari Ollikainen, David
484Pullinger, John Smith, Edward Vielmetti, and Jane
485Williams. The useful role which NCSA's XMosaic
486information browser tool played in assembling the
487information on which this report was based should also be
488acknowledged - many thanks to its developers.
489
490All trademarks are hereby acknowledged as being the
491property of their respective owners.
492
493 Disclaimer
494
495This report is based on information supplied to or
496obtained by Edinburgh University Computing Service (EUCS)
497in good faith. Neither EUCS nor RARE nor any of their
498staff may be held liable for any inaccuracies or
499omissions, or any loss or damage arising from or out of
500the use of this report.
501
502The opinions expressed in this report are personal
503opinions of the author. They do not necessarily
504represent the policy either of RARE or of ECUS.
505
506Mention of a product in this report does not constitute
507endorsement either by EUCS or by RARE.
508
509 Availability
510
511This document is available in various forms (PostScript,
512text, Microsoft Word for Windows 2) by anonymous FTP
513through the following URL:
514
515 ftp://ftp.edinburgh.ac.uk/pub/mmaccess/
516
517Paper copies may be available from RARE.
518
519 Introduction
520
521 Background
522
523This study was inspired by the realisation that while
524some aspects of distributed multimedia technology are
525being actively introduced into the European research
526community (for instance, audiovisual conferencing,
527through the MICE project), other aspects are receiving
528less attention. In particular, one category in which
529there seems to be relatively little activity is providing
530solutions to ease remote access to multimedia resources
531(for instance, accessing stored audio/video clips or
532images, or indeed entire multimedia applications, across
533the network). Few commercial products address this, and
534the relevance of existing standards in this area is
535unclear.
536
537Of the 50 or so research projects documented in the
538recent RARE distributed multimedia survey [Adi93], only
539about six have a direct relevance to this application
540area. Where stated in the survey, the main research
541effort in these projects is often directed towards the
542"difficult" problems, such as the transfer of isochronous
543data and the design and implementation of object-oriented
544multimedia databases, rather than towards user-oriented
545issues.
546
547This report is concerned with practical issues in the
548intersection of networked information retrieval, database
549and multimedia technologies. It aims to establish actual
550user requirements in this area, to look at existing
551systems which offer partial solutions, and to identify
552what additional work needs to be done to satisfy the most
553pressing requirements.
554
555 Terminology
556
557In order to discuss multimedia information systems, we
558need a consistent terminology. The vocabulary defined
559below embodies some of the concepts of the Dexter
560hypertext reference model [HS90]. This model is
561sufficiently general to be useful for describing most of
562the facilities and requirements of the multimedia
563information systems described in this report. (However,
564the Dexter model does not describe searchable index
565objects - it is not a database reference model.)
566
567anchor An identified portion of a node. Eg in a
568 text node, an anchor might be a string of
569 one or more adjacent characters, while in
570 an image node it might be a rectangular
571 area of the image.
572
573composite node A node containing data of multiple media
574 types.
575
576document Often used loosely as a synonym for node.
577
578hyperdocument We refer to a collection of related nodes,
579 linked internally with hyperlinks, as a
580 "hyperdocument". Examples are a database
581 of medical images and associated text; a
582 module from a suite of teaching material;
583 or an article in a scientific journal. A
584 hyperdocument may contain hyperlinks to
585 other data which exists in other
586 hyperdocuments, but can be viewed as
587 largely self-contained. It is a high-
588 level "unit of authoring", but is not
589 necessarily perceived as a distinct unit
590 by a reader (although it may be so
591 perceived, particularly if it contains few
592 hyperlinks to outside entities).
593
594hyperlink Set of one or more source anchors and one
595 or more target anchors. Also known simply
596 as a "link".
597
598isochronous (adjective) Describes a continuous flow of
599 data which is required to be delivered by
600 the network under critical time
601 constraints.
602
603leaf node A node which contains no source anchors.
604
605media type An attribute of data which describes the
606 general nature of its expected
607 presentation. The value of this attribute
608 could be one of the following (not
609 exhaustive) list:
610
611 Text
612 Sound
613 Image (eg a "photograph")
614 Graphics (eg a "drawing")
615 Animation (ie moving graphics)
616 Movie (ie moving image)
617
618monomedia (adjective) Said of data which is all of the
619 same media type.
620
621multimedia (adjective) Said of data which contains
622 different media types. This definition is
623 stricter than general usage, where
624 "multimedia" is often used as a generic
625 term for non-textual data, and where it
626 may even be used as a noun.
627
628physical media Magnetic or optical storage. Not to be
629 confused with media type!
630
631[simple] node A monomedia object which may be retrieved
632 and displayed as a single unit.
633
634source anchor An anchor which may be "actioned" by the
635 user, causing the node(s) containing the
636 target anchor(s) in the same hyperlink to
637 be retrieved and displayed. This process
638 is called "traversing the link".
639
640target anchor an anchor forming part of a hyperlink,
641 whose containing node is retrieved and
642 displayed when the hyperlink is traversed.
643
644 User Requirements
645
646User requirements in an area such as networking, which is
647subject to rapid technological change, are sometimes
648difficult to identify. To an extent, technology leads
649applications, and users will exploit what is possible.
650
651 Applications
652
653Awareness of the range of networked multimedia
654applications which are currently being envisaged by
655computer users in the academic and research community
656leads to a better understanding of the technical
657requirements. This section outlines some projects which
658require remote access to multimedia information across
659research networks, and which are currently either at a
660preliminary stage or underway. The projects are divided
661into broad categories according to their characteristics.
662
663 Multimedia Databases
664
665Here are several examples of multimedia projects which
666have a "database" character.
667
668The Peirce Telecommunity Project
669
670 This project centres on the construction of a
671 multimedia (text and image) database of the works of
672 the American philosopher Peirce, together with tools
673 to process the data and to make it available over
674 the Internet. A sub-project at Brown University
675 focuses on adapting existing client/server network
676 tools for this purpose. The requirements for
677 network access include facilities for structured
678 viewing, intelligent retrieval, navigation, linking,
679 and annotation, as well as for domain-specific
680 processing.
681
682Museum Object Databases
683
684 The RAMA (Remote Access to Museum Archives) project
685 is funded under the EEC RACE II programme. Its
686 objective is to develop a system which allows
687 museums to make multimedia information about their
688 exhibits and archived material available over an
689 ISDN network. The requirements capture and
690 technical architecture design phases are now
691 complete, and a prototype system will be delivered
692 in June 1993 to link the Ashmolean Museum (Oxford,
693 GB), the Musee d'Orsay (Paris, FR) and the Museum
694 Archeological National (Madrid, ES). Image data is
695 the main media type of interest, although video and
696 sound may also play a part.
697
698The Bristol Biomedical Videodisk Project
699
700 The Bristol Biomedical Videodisc is a collection of
701 Medical, Veterinary and Dental images. The
702 collection holds some 24,000 still images and is
703 continuously growing. Textual information regarding
704 the images is included as part of the database and
705 this can be searched on any keyword, number or other
706 data type, or a combination of any of these. The
707 images are currently delivered in analogue form on a
708 videodisc, but many institutions are unable to
709 afford the cost of videodisc players.
710 Investigations into making this image and text
711 database available across the network are underway.
712
713ArchiGopher
714
715 ArchiGopher is a Gopher server at the College of
716 Architecture, University of Michigan, dedicated to
717 the dissemination of architectural knowledge.
718 Presently in its infancy, ArchiGopher is intended to
719 become a multimedia resource for all architecture
720 faculty and students world-wide. Some of the
721 available or planned resources are:
722
723 The College's image bank.
724 The CAD group's collection of computer models
725 (already started).
726 The Doctoral Program's recent dissertation
727 proposals and abstracts.
728 Example archive of Kandinsky paintings.
729 Images of 3D CAD projects
730
731 The principal media type in ArchiGopher is image.
732 Files are stored in both TIFF and GIF format.
733
734Vatican Library Exhibit
735
736 In January 1993, the US Library of Congress mounted
737 an electronic version of the exhibition ROME REBORN:
738 THE VATICAN LIBRARY AND RENAISSANCE CULTURE. The
739 exhibition was subsequently processed by the
740 University of Virginia Library. The text files were
741 broken into individual captions associated directly
742 with each image and a WAIS-searchable version of the
743 object index generated. This has been made
744 available on Gopher by the University of Virginia
745 Library.
746
747 This project is particularly interesting, as it
748 demonstrates some limitations of the Gopher system.
749 The principal media types are image and text, and it
750 is difficult to associate a caption with its image -
751 each must be fetched separately, and using the
752 XMosaic or xgopher client software it is not
753 possible to tell which menu entry is the image and
754 which the caption. (This may be a consequence of
755 how the data has been configured for the Gopher
756 server; if so, a requirement for better publishing
757 tools may be indicated.) Furthermore, searching the
758 object index will result in a Gopher menu containing
759 references to catalogue entries for relevant
760 exhibits, but not to the on-line images of the
761 exhibits themselves, which severely limits the
762 usefulness of the index.
763
764 It is interesting to note that during the
765 preparation of this report, the Vatican Exhibition
766 has been mounted on the World-Wide Web (WWW). The
767 hypermedia presentation on the Web is very much more
768 attractive to use than the Gopher version.
769
770Jukebox
771
772 Jukebox is a project supported by the EEC libraries
773 program. The project aims to evaluate a pilot
774 service providing library users with on-line access
775 to a database of digital sound recordings. The
776 database will support multi-user access and use
777 suitable storage media to make available sound
778 recordings in a compressed format. Users will
779 access the service with a personal computer
780 connected to a telematic network.
781
782 Scientific Publishing
783
784There are several refereed electronic academic journals
785presently distributed on the Internet. These tend to be
786text-only journals, and have not really addressed the
787issues of delivering and manipulating non-text data.
788
789Many scientific publishers have plans for electronic
790publishing of existing academic journals and conference
791proceedings, either on physical media or on the network.
792The Journal of Biological Chemistry is now published on
793CD-ROM, for instance. Some publishers view CD-ROM as an
794interim step to the ultimate goal of making journals
795available on-line on the Internet.
796
797The main types of non-text data which are envisaged are:
798
799 Images. In many cases, image data (a microphotograph,
800 say) is central to an article. Software which
801 recognises that the text may be of secondary importance
802 to the image is required.
803
804 Application-specific data. The ChemLab and MoleculeLab
805 applications are widely used, and the integration of
806 corresponding data types with journal articles will
807 enhance readers' ability to visualise molecular
808 structures. Similarly, mathematics appearing in
809 scientific papers could be represented in a form
810 suitable for processing by applications such as
811 Mathematica. Mathematical content could then become a
812 much more interactive and dynamic aspect of research
813 publications.
814
815 Tabular data. The ability for a reader to extract
816 tabular data from a research paper, to produce a
817 graphical representation, to subset the data, and to
818 further process it in a number of different ways, is
819 viewed as an essential part of scientific electronic
820 publishing.
821
822 Movies. The American Astronomical Society regularly
823 publishes videos to go with its academic journals.
824 Electronic publishing can improve on this "hard copy"
825 publishing by integrating video data much more closely
826 with the source article.
827
828 Sound. There is perhaps slightly less demand for audio
829 information in scientific publishing, but the
830 requirement does exist in particular specialities (such
831 as acoustics and zoology journals).
832
833Access to academic journals using at least four different
834paradigms is envisaged. Hierarchical access, perhaps
835using a traditional journal/volume/issue/article model,
836is perhaps the most obvious. Keyword searching (or full-
837text indexing) will be required. Browsing is another
838useful and often underestimated access model - to support
839browsing it is essential that "eye-catching" data
840(unlikely to be textual) is prominently accessible. The
841final method of access is perhaps the most important -
842the use of interactive viewing tools. Such tools would
843enable navigation of hypermedia links within and between
844articles, with gateways to special-purpose applications
845as described above. The use of these disparate access
846methods implies more than one structure being applied to
847the same underlying data.
848
849Standards, particularly SGML, are becoming important to
850publishers, and it is clear that the SGML-based HyTime
851standard will be a front runner in providing the kind of
852hypermedia facilities which are being envisaged.
853However, progress towards a common SGML Document Type
854Definition (DTD) for scientific articles, even within
855individual publishing houses and for text-only documents,
856is slow.
857
858A specific initiative involving interested parties will
859be required to formalise detailed requirements and to
860pilot standards in this area. A preliminary demonstrator
861project, funded by publishers and by the British Library
862Research and Development Department, involves making
863about 30 sample scientific articles available over the
864SuperJANET network, using a range of different software
865products. The demonstrator project is being managed by
866IOP Publishing and is being carried out at Edinburgh
867University Computing Service.
868
869Existing tools, particularly WAIS and WWW, are relevant,
870but adequate security and charging mechanisms are
871required if commercial publishers are to use them. Many
872research groups are now making the text of preprints and
873published research papers available on Gopher servers.
874
875It is interesting to note that the proceedings of the
876Multimedia 93 conference run by the ACM will be published
877electronically (on CD-ROM), using a multimedia document
878format designed specifically for the event.
879
880 Computer-aided Learning
881
882The ready availability of user-friendly multimedia
883authoring tools such as AuthorWare Professional,
884Asymmetrix Multimedia Toolbook, Macromind Director and
885many more, has stimulated much interest in multimedia for
886computer-aided learning applications within the user
887community. Sophisticated interactive multimedia
888courseware applications are being developed in many
889disparate subjects throughout the European academic
890community. Users are now beginning to ask network
891technologists, "how can I make my multimedia application
892available to others across the network?".
893
894There is considerable interest in using the network to
895enhance delivery of multimedia teaching materials - for
896instance to allow students to take courses remotely
897(distance learning) and for their learning process to be
898supported, monitored and assessed remotely.
899
900The requirements which flow from this type of network
901application include the ability to identify and
902authenticate the students using the material, to monitor
903their progress, and to supply on-line assessment
904exercises for the student to complete. Multimedia
905authoring tools allow very attractive presentation
906environments to be created, which encourages learning;
907this is viewed as essential by course developers. Easy-
908to-use authoring tools (preferably existing commercial
909ones) are also essential.
910
911Finally, some learning applications involve simulations -
912examples include meteorological modelling and economic
913simulations. Network delivery of teaching materials
914should cope with this requirement (perhaps by
915acknowledging that executable scripts are just another
916media type).
917
918 General Information Services
919
920There are many other possible uses of multimedia data in
921networked information servers which don't conveniently
922fall into any of the above categories. Some examples are
923given below.
924
925 On-line documentation. Manuals and instruction books
926 often rely heavily on pictorial information, and are
927 enhanced by dynamic media types (sound, video). The
928 ability to access centrally-held manuals across a
929 network makes it much easier to keep the information up-
930 to-date.
931
932 Campus-wide information systems (CWIS) are an important
933 growth area. The opportunities for enhancing such a
934 service with multimedia data (eg maps) is obvious.
935
936 Multimedia news bulletins (eg the Internet Talk Radio,
937 which is sound only).
938
939 Product information (the multimedia equivalent of paper
940 advertising matter).
941
942 Consumer systems - eg tourist information servers. The
943 utility of such systems in an academic/research
944 environment is perhaps questionable, but it is likely
945 that such systems will address problems which will also
946 be met in this environment. We should be prepared to
947 learn from such projects.
948
949 Data Characteristics
950
951Some of the characteristics which make data more
952appropriate for network publication rather than
953publication on physical media are listed below.
954
955 The data may change frequently.
956
957 Implementing corrections and improvements to the data
958 is very much easier.
959
960 It is more readily available to the data user - no
961 purchase/delivery cycle need exist.
962
963 Publication on physical media may not be cost-effective
964 for very large volumes of data. (Of course, there is a
965 cost in networking the data as well, but the
966 research/academic user is normally insulated from
967 this.)
968
969 Access for large user communities can be established
970 without requiring each user to purchase a potentially
971 expensive physical media peripheral (such as a laser
972 disk player). This is particularly helpful in
973 classroom situations.
974
975 It may require less effort from the data publisher to
976 make data available over a network, rather than set up
977 a manual mechanism for distributing physical media.
978
979 If related data from many different sources is to be
980 published, it may be more efficient to leave the data
981 in situ, and simply publish the network addresses of
982 the data.
983
984There are counter-reasons which may make physical media
985distribution more appropriate:
986
987 Easier to charge for. (However, charging mechanisms do
988 exist in some network information systems. It may be
989 that potential information providers need to be made
990 more aware of this.)
991
992 Easier to deter or prevent copyright infringement,
993 using traditional copy-protection techniques.
994
995 Requirements Definition
996
997From studying the applications described in the preceding
998section, and from discussions with the people involved
999with the applications, it is possible to draw up a list
1000of general requirements which a distributed multimedia
1001information system for the academic and research
1002community should satisfy. These requirements are
1003informally described in the following subsections. The
1004descriptions are necessarily informal and incomplete:
1005every individual application will have its own detailed
1006requirements, which would take a great deal of effort to
1007determine (and indeed some of the requirements may not
1008become apparent until the application is into its
1009development phase).
1010
1011 Platforms
1012
1013It is clear that the European academic community, in
1014common with other such communities, requires support for
1015three main platforms: UNIX, Apple Macintosh, and
1016PC/Windows. For multimedia client/server systems, the
1017latter two are less appropriate as server platforms, but
1018client support for all three is vital. UNIX will be most
1019often used as the server platform.
1020
1021There are other systems, such as VAX/VMS, which are also
1022important in some sectors.
1023
1024 Media Types
1025
1026Unsurprisingly, all applications require text data to be
1027supported as a basic media type. Image and graphic media
1028types are next in importance, followed by "application-
1029specific" data (such as tabular scientific data,
1030mathematical equations, chemical data types, etc). Sound
1031and video media types are becoming more important as
1032users discover how these can enhance applications.
1033
1034Many different encodings are possible for each media type
1035(eg image data can be encoded as TIFF, PCX, GIF, PICT and
1036many more). An information system should not constrain
1037the type of encoding used, and should ideally offer
1038either a range of alternative encodings, or conversion
1039facilities between the stored encoding and an encoding
1040suitable for display by the client workstation.
1041
1042 Hyperlinks
1043
1044It is clear that many applications require their users to
1045be able to navigate through the information base
1046according to relationships determined by the information
1047provider - in other words, hyperlinks. Academic
1048publishing, CAL, on-line documentation and CWIS systems
1049all require this capability. The user should be able, by
1050some action such as clicking on a highlighted word in a
1051text node or on a button, to cause another node or nodes
1052to be retrieved and displayed.
1053
1054Some "hypermedia" systems are in fact simply hypertext,
1055in that they require the source anchor of a hyperlink to
1056be in a text node. A true hypermedia system allows
1057hyperlinks to have their source anchors in nodes of any
1058media type. This allows a user to click the mouse on a
1059component of a diagram or on part of a video sequence to
1060cause one or more related nodes to be retrieved and
1061displayed.
1062
1063Some hypermedia systems allow target anchors of a
1064hyperlinks to be finer-grained than a whole node - eg the
1065target anchor could be a word or a paragraph within a
1066text document. Without such a capability, it is
1067necessary for target nodes to be quite small if precision
1068is required in a hyperlink. This may be difficult to
1069manage, and fine-grained target anchors are therefore
1070better.
1071
1072Additional structure above or orthogonal to the
1073underlying hyperlinked data is required in some
1074applications. This allows the same (generally non-
1075textual) data to be used in several different
1076applications, or the implementation of different access
1077paradigms.
1078
1079 Presentation
1080
1081Related information of different media types must be
1082capable of synchronised display. Commercial multimedia
1083authoring packages provide many different ways of
1084presenting, synchronising and interacting with media
1085elements. Some of these are summarised below.
1086
1087 Backdrops. An application may present all its visual
1088 information against a single background bitmap - eg a
1089 CAL application might use a background image of an open
1090 textbook, with graphics, text and video data all
1091 presented on the open pages of the book.
1092
1093 Buttons. A "button" can be defined as an explicitly-
1094 delimited area of the display, within which a mouse
1095 click will cause an action to occur. Typically, the
1096 action will be (or can be modelled as) a hyperlink
1097 traversal. Applications use different styles of button
1098 - some may use "tabs" as in a notebook, or perhaps
1099 "bookmarks" in conjunction with the open textbook
1100 backdrop mentioned above. Others may use plain buttons
1101 in a style conforming to the conventions of the host
1102 platform, or may simply highlight a word or phrase in a
1103 text display to indicate it is "active".
1104
1105 Synchronisation in space. When two or more nodes are
1106 presented together (eg because a link with more than
1107 one target anchor has been traversed), the author of
1108 the hyperdocument may wish to specify that they be
1109 presented in a spatially-related way. This may
1110 involve: x/y synchronisation - eg a video node being
1111 displayed immediately above its text caption; it may
1112 involve contextual synchronisation - eg an image being
1113 displayed in a specific location within a text node; or
1114 it may involve z-axis synchronisation as well - for
1115 instance a text node containing a simple title being
1116 displayed on top of an image, with the text background
1117 being transparent so that the image shows through.
1118
1119 Synchronisation in time. Isochronous data may require
1120 synchronisation - the obvious case being audio and
1121 video tracks (where these are held separately). Other
1122 examples are: the synchronisation of an automatically-
1123 scrolling text panel to a video clip (for subtitling);
1124 or to an audio clip (eg a translation); or
1125 synchronising an animation to an explanatory audio
1126 track.
1127
1128 Searching
1129
1130Database-type applications require varying degrees of
1131sophistication in retrieval techniques. For applications
1132addressed in this report, non-text nodes form the major
1133data of interest. Such nodes have associated
1134descriptions, which may be plain text, or may be
1135structured into fields. Users need to be able to search
1136the descriptions, obtain a list of "hits", and select
1137nodes from that list to display. Searching requirements
1138vary from simple keyword searching, via full-text
1139indexing (with or without Boolean combinations of search
1140words), to full SQL-style database retrieval languages.
1141
1142 Interaction
1143
1144The user must be able to annotate documents retrieved
1145from the information server. The annotations may be
1146stored locally. Similarly, the user may wish to add his
1147own (locally-held) hyperlinks to documents. (Actual
1148modification of documents in the information system
1149itself, or shared annotations to documents - ie the
1150information system as a CSCW environment - is viewed as
1151separate issue which this report does not address.)
1152
1153If an information provider has included contact details
1154(such as a mail address) in a document, it should be
1155possible for the reader to invoke a program (such as a
1156mailer) which initiates communication with the author.
1157
1158In some applications, it may make sense for a user to be
1159able to specify a region of interest in an image or movie
1160clip, and to request a more detailed view of (or other
1161information about) that region.
1162
1163Some applications require a sequence of images to be
1164presented under control of the user. For instance, a
1165three-dimensional microscopic structure could be
1166represented as a sequence of images taken with the
1167microscope focused on a different plane for each image.
1168For display, the user could control which image was
1169displayed using some kind of slider control, giving the
1170illusion of focusing a microscope. (This particular
1171example has been taken from the Theseus project at John
1172Moore's University, Liverpool, GB.)
1173
1174 Quality of Service
1175
1176Research has shown [Shn84] that user toleration of delay
1177in computer systems depends on user perception of the
1178nature of the requested action. If the user believes
1179that no computation is required, tolerable delays are of
1180the order of 0.2s. If the user believes the action he or
1181she has requested the computer to perform is "difficult"
1182- for instance a computation of some form - then a
1183tolerable delay is of the order of 2s. Users tend to
1184give up waiting for a response after about 20s.
1185Networked multimedia information systems must be able to
1186provide this level of responsiveness.
1187
1188 Management
1189
1190In order to support applications involving real-money
1191information services (eg academic publishing) and
1192learning/assessment applications, there must be a
1193reliable and secure access control mechanism. A simple
1194password is unlikely to suffice - Kerberos authentication
1195procedures are a possibility.
1196
1197Users must be able to determine the charge for an item
1198before retrieving it (assuming that pay-per-item will be
1199a common paradigm - alternatives such as pay-per-call,
1200pay-per-duration are also possible). Access records must
1201be kept by the information server for charging purposes.
1202
1203Learning applications have similar requirements, except
1204that the purpose here is not to charge for information
1205retrieved, but to monitor and perhaps assess a student's
1206progress.
1207
1208 Scripting
1209
1210Many authoring packages provide scripting languages. In
1211most cases, these languages are used to manage the
1212presentation environment and control navigation within
1213the hypermedia document. There are other, declarative
1214rather than procedural, methods for achieving this, so
1215scripting of this type is not necessarily a requirement.
1216However, some application areas require executable
1217scripts for other purposes (eg simulations in CAL
1218applications). Care in providing such a facility is
1219required, because of the potential for abuse (the
1220possibility of "trojan" scripts). However, there is work
1221going on to produce "safe" scripting languages - an
1222example is "safe tcl", being developed by Borenstein and
1223Ousterhout (contact ouster@cs.berkeley.edu).
1224
1225 Bytestream Format
1226
1227For the easy transfer and handling of a hyperdocument, it
1228must be capable of being encoded into a bytestream form,
1229in such a way that the structure of the document is
1230preserved and it can be decoded without loss of
1231information.
1232
1233This facility makes it possible for such documents to be
1234supplied to a user over electronic mail, in such a way
1235that he or she can browse them at his or her own site.
1236This may be appropriate where the user does not have a
1237direct connection to the Internet. It will also be
1238useful for printing the hyperdocument.
1239
1240 Authoring
1241
1242It is essential that a multimedia information system
1243should have adequate authoring tools which make it easy
1244to prepare and publish hypermedia information. Such
1245tools need similar power to existing commercial
1246multimedia authoring software for stand-alone multimedia
1247applications.
1248
1249 Existing Systems
1250
1251This chapter describes some existing distributed
1252information systems in sufficient detail to reveal how
1253they handle multimedia data, and analyses how well they
1254meet the requirements outlined in the preceding chapter.
1255
1256 Gopher
1257
1258The Internet Gopher is a distributed document delivery
1259service. It allows a neophyte user to access various
1260types of data residing on multiple hosts in a seamless
1261fashion. This is accomplished by presenting the user
1262with a hierarchical arrangement of nodes and by using a
1263client-server communications model. The Gopher server
1264accepts simple queries, and responds by sending the
1265client a node (usually called a document in this
1266context).
1267
1268Client software is available for a large number of
1269systems, including:
1270
1271 UNIX (character terminals)
1272 X windows
1273 Apple Macintosh
1274 MS DOS
1275 NeXT
1276 VM/CMS
1277 VMS
1278 OS/2
1279 MVS/XA
1280
1281Servers are available for systems such as:
1282
1283 UNIX
1284 VMS
1285 Apple Macintosh
1286 VM/CMS
1287 MVS
1288 MS DOS
1289
1290Gopher was developed at the University of Minnesota.
1291
1292 Gopher User Image
1293
1294A Gopher client offers an interface into "gopherspace",
1295which appears to the user as a hierarchy of menus and
1296document nodes, similar in some ways to a file system
1297hierarchy of directories and files. Selecting an entry
1298from a menu node causes a further menu to appear, or
1299causes a document to be retrieved and displayed.
1300
1301As well as "ordinary" document nodes, Gopher has "search
1302nodes" - when one of these is selected from a menu, the
1303user is prompted for one or more words to search on. The
1304result of the search is a "virtual" menu, containing
1305entries for document nodes (within some subset of
1306gopherspace) which match the search. A special type of
1307Gopher search server called "veronica" provides access to
1308a database of all directory nodes in gopherspace. This
1309allows a user to construct a virtual menu of all Gopher
1310menu items containing a particular word. WAIS databases
1311may also be located at Gopher search nodes, since some
1312Gopher servers understand the format of WAIS index files.
1313
1314 Gopher Protocol
1315
1316Gopher uses a client-server paradigm. The Gopher
1317protocol runs over a reliable data stream service,
1318typically TCP, and is fully defined in RFC 1436. The
1319following paragraphs give an overview which is sufficient
1320for understanding how multimedia data is handled in
1321Gopher.
1322
1323A Gopher client opens a TCP connection to a Gopher server
1324(defined by machine name and TCP port number), and sends
1325a line of text known as the "selector" to request
1326information from the server. The server responds with a
1327block of data, and then closes the connection. No state
1328is retained by the server. A null (empty) selector tells
1329the Gopher server to return its "root" menu node,
1330containing pointers to other information in gopherspace.
1331
1332A menu is returned from a Gopher server as a sequence of
1333lines of text, each corresponding to one entry in the
1334menu. Each line (which is sometimes called a "Gopher
1335reference") contains the following data, which can be
1336used by the client software to retrieve and display the
1337corresponding node in gopherspace.
1338
1339 A single character which identifies the type of the
1340 node. Possible values of this type ID are given below.
1341 A human-readable string which is used by the client
1342 software when it displays the menu entry to the user.
1343 The selector which should be used by client software to
1344 retrieve the node. It is treated as opaque by the
1345 client software.
1346 The domain name of the host on which the node is held.
1347 The port number to use for the TCP connection.
1348
1349A document node is sent by a Gopher server simply as
1350lines of text terminated by a dot on a line by itself, or
1351as raw binary data, with the end of the data indicated by
1352the server closing the TCP connection. The choice
1353depends on the type of node.
1354
1355The currently-defined type IDs are as follows:
1356
13570 Node is a file.
13581 Node is a directory.
13592 Node is a CSO phone book server.
13603 Error.
13614 Node is a BinHexed Macintosh file.
13625 Node is DOS binary archive of some sort.
13636 Node is a UNIX uuencoded file.
13647 Node is a search server.
13658 Node points to a text-based telnet session.
13669 Node is a binary file.
1367T Node points to a TN3270 connection.
1368
1369Some experimental IDs are also in use:
1370
1371s Node contains -law sound data.
1372g Node contains GIF data.
1373M Node contains MIME data.
1374h Node contains HTML data.
1375I Node contains image data of some kind.
1376i In-line text type.
1377
1378The process for defining new data types and corresponding
1379IDs is not clear.
1380
1381 Gopher+ Protocol
1382
1383The Gopher+ protocol is an extension of the Gopher
1384protocol. Gopher+ is defined informally in [AAL92]. It
1385is designed to be downwards compatible with the original
1386protocol, so that old Gopher clients may access Gopher+
1387servers (without being able to take advantage of the new
1388facilities), and Gopher+ clients may access old Gopher
1389servers. Gopher+ is still at the experimental stage, and
1390is liable to change.
1391
1392The most important new feature is the introduction of
1393"attributes" associated with individual nodes. The
1394client may retrieve the attributes of a node instead of
1395the node contents. Attributes defined so far include:
1396
1397INFO Contains the Gopher reference of the node.
1398 Mandatory.
1399
1400ADMIN Contains administrative information,
1401 including the mail address of the server
1402 administrator and the last-modified date
1403 of the node. Mandatory.
1404
1405VIEWS Contains a list of one or more "view
1406 descriptors", each of which describes an
1407 alternate view of the node. For instance,
1408 an image node may contain a TIFF view, a
1409 GIF view, a JPEG view, etc. The client
1410 software (or the user) may choose which
1411 view to retrieve. The size of the view is
1412 also (optionally) available in this
1413 attribute. The Gopher+ Attribute Registry
1414 (see below) defines the permitted view
1415 types.
1416
1417ABSTRACT This attribute contains a short
1418 description of the item. It may also
1419 include a Gopher reference to a longer
1420 abstract, held in a separate Gopher node.
1421
1422ASK This attribute is used for the interactive
1423 query extension. The interactive query
1424 facility in Gopher+ is used to obtain
1425 information from a user before retrieving
1426 the contents of a node. The client
1427 fetches the ASK attribute, which contains
1428 a list of questions for the user. His or
1429 her responses to those questions are sent
1430 along with the selector to the server,
1431 which then returns the contents of the
1432 node. This facility could be used as a
1433 very simple way of querying a database,
1434 for instance. Using the interactive query
1435 facility to supply a password for access
1436 control purposes is not a good idea -
1437 there are too many opportunities for
1438 masquerading.
1439
1440The University of Minnesota maintains a registry of
1441Gopher+ attribute types. For the VIEWS attribute, the
1442registry contains a list of permitted view types. Note
1443that these view types have a similar function to the type
1444identifier described in the preceding section.
1445
1446The general format of a Gopher+ view descriptor is:
1447
1448 xxx/yyy zzz: <nnnK>
1449
1450where xxx is a general type-of-information advisory, yyy
1451is what information format you need understand to
1452interpret this information, zzz is a language advisory
1453(coded using POSIX definitions), and nnn is the
1454approximate size in bytes. Possible values for xxx
1455include text, file, image, audio, video, terminal.
1456
1457(It now appears that the University of Minnesota Gopher
1458Team accepts the need to be consistent in the use of
1459type/encoding attributes with the MIME specification.
1460The Gopher+ Type Registry may thus eventually disappear,
1461together with the set of xxx/yyy values it currently
1462contains.)
1463
1464No view descriptors for directory nodes are currently
1465registered.
1466
1467In order to make use of the information available in
1468attributes, it is necessary to fetch the attributes
1469before fetching the contents of a node. Gopher+ provides
1470a way of fetching the attributes for each entry in a menu
1471at the same time as the menu is retrieved. This saves
1472having to establish two successive TCP connections to
1473fetch a single document, at the expense of some
1474additional client software complexity.
1475
1476 Gopher Publishing
1477
1478The procedure for making data available using the Unix
1479Gopher server "gopherd" is very straightforward. The
1480hierarchical nature of the Unix file system closely
1481matches the Gopher concept of menus and documents. The
1482gopherd program exploits this - Unix directories are
1483represented as Gopher menu nodes, and Unix files as
1484Gopher document nodes. The names of directories and
1485files are the entries in Gopher menus. This can lead to
1486awkward file names containing spaces, so gopherd provides
1487an aliasing mechanism (the .cap directory) to get round
1488this.
1489
1490To represent menu entries pointing to Gopher nodes on
1491other servers, special "link" files (starting with a dot)
1492are used.
1493
1494The type ID for a document node is determined from the
1495extension of its Unix filename. If a client requests a
1496file containing a shell script, the script is executed
1497and the output returned to the client.
1498
1499The Gopher+ version of gopherd is similar, but the .cap
1500directory is replaced by a configuration file
1501gopherd.conf. This file is used to specify
1502administration attributes, and the mapping between
1503filename extensions and view descriptors. Some limited
1504access control (based on the client's IP address/domain
1505name) is also provided by the Gopher+ version of gopherd.
1506
1507 Published Non-text Data
1508
1509There is already some useful non-text data published on
1510Gopher - almost exclusively image data. See for example
1511the Vatican Library Exhibition at the University of
1512Virginia Library, the ArchiGopher at the University of
1513Michigan, the weather machine at the University of
1514Illinois. Some of these are described in the User
1515Requirements chapter of this report.
1516
1517There seem to be rather fewer sound archives in
1518gopherspace, but interested users may access the
1519Edinburgh University Computing Service Gopher server on
1520gopher.ed.ac.uk, where the Testing Area contains 20 or 30
1521short audio files in Sun audio format. Note - the
1522availability of this archive is not guaranteed.
1523
1524 Advantages
1525
1526The main factor in favour of Gopher is its widespread
1527penetration. There are over 1000 Gopher servers world-
1528wide. This popularity is due in part to the ease of
1529setting up a Gopher server and making information
1530available on it, particularly on a Unix platform.
1531
1532 Limitations
1533
1534It is unfortunate that the relatively well-defined MIME
1535types were not adopted in Gopher+. As mentioned above,
1536this may yet happen, although there appear to be reasons
1537for keeping the set of MIME types small whereas Gopher
1538requires a wide range of types to offer to clients. The
1539latest word is that the MIME registry will be expanded to
1540include the types which the Gopher+ developers want.
1541
1542Gopher is inflexibly hierarchical in nature. Hypertext
1543or hypermedia it is not - links to other nodes from
1544within document nodes are not possible. There is a
1545suggestion in the Gopher+ specification that alternate
1546views of directory nodes could be used to provide some
1547kind of hypermedia capability, but this does not yet
1548exist, and it is unlikely that it could be made to work
1549as easily as the WWW hypertext model.
1550
1551There is no access control at the user level - anyone can
1552retrieve anything on a Gopher server. There is no
1553provision for charging for information.
1554
1555 Wide Area Information Server
1556
1557The Wide Area Information Server (WAIS) system allows
1558users to search for and retrieve information from
1559databases anywhere on the Internet. WAIS uses a client-
1560server paradigm, and client and server software is
1561available for a wide range of platforms. Client
1562applications are able to retrieve text or other media
1563documents stored on the servers, by specifying keywords.
1564The server software searches a full-text index of the
1565documents, and returns a list of documents containing the
1566keywords (ranked according to a heuristic algorithm).
1567The client may then request the server to send a copy of
1568any of the documents found. Relevant documents can be
1569fed back to a server to refine the search. Successful
1570searches can be automatically re-run, to alert the user
1571when new information becomes available.
1572
1573WAIS was developed by Thinking Machines Corporation of
1574Cambridge, Massachusetts, in collaboration with Apple
1575Computer Inc., Dow Jones and company, and KPMG Peat
1576Marwick. The WAIS software has been made freely
1577available; however Thinking Machines has announced that
1578they will stop support for their publicly-distributed
1579WAIS as of version 8-b5.1. Future support and
1580development of the publicly-distributed WAIS has been
1581taken over by CNIDR (Clearinghouse for Networked
1582Information Discovery and Retrieval) in the USA. Future
1583CNIDR releases will be called FreeWAIS. A new company,
1584WAIS Inc, has been formed by Thinking Machines to take
1585over commercial exploitation of the Thinking Machines
1586WAIS software.
1587
1588WAIS server software is available for the following
1589platforms:
1590
1591 UNIX
1592 VAX/VMS
1593
1594Client software is available for the following platforms:
1595
1596 UNIX (versions for X, Motif, Open Look, Sun View)
1597 NeXT
1598 Macintosh
1599 MS DOS
1600 MS Windows
1601 VAX/VMS
1602
1603There are currently over 400 WAIS databases available on
1604the Internet. WAIS is also the basis of some commercial
1605information services on private networks.
1606
1607 WAIS User Image
1608
1609In order to ask a question, the user must first select
1610one or more databases in which to look for the answer.
1611(The list of all available databases is available from a
1612number of well-known sites.) The next step is to enter
1613one or more keywords as the basis of the search. The
1614search will return a list of documents (the "result set")
1615which contain any of the keywords. Each document is
1616given a ranking (a number between 1 and 1000) which
1617indicates how relevant to the user's question the server
1618believes the document to be. The size of each document
1619is also shown in the list. The user may limit the size
1620of the result set - the default limit is typically 40
1621documents.
1622
1623The user may then choose to retrieve and display one or
1624more documents from the list. Alternatively, he or she
1625may designate one or more documents in the list as
1626"relevant", and perform another search to find "more
1627documents like this". This is called "relevance
1628feedback".
1629
1630The user may retrieve general information about the
1631database, and may examine the catalogue of all documents
1632in the database. There is also a "database of
1633databases", which may be searched to identify WAIS
1634databases which may be relevant to a subject.
1635
1636 WAIS Protocol
1637
1638The user interface (client) talks to the server using an
1639extended version of a standard ANSI protocol called
1640Z39.50. This is now aligned with the ISO SR (Search and
1641Retrieval) protocol for bibliographic (library)
1642applications, which is part of OSI. The present WAIS
1643protocol does not utilise a full OSI stack - APDUs are
1644transferred directly over a TCP/IP connection. The WAIS
1645protocol is described in [DKM90].
1646
1647WAIS does not, at this time, implement the full Z39.50-
16481992 specification - in particular, WAIS does not permit
1649Boolean searches (eg "find all documents containing
1650'chalk' and 'cheese' but not 'green'"). However, Boolean
1651search capability is being added to the FreeWAIS
1652implementation. There are facilities in the Z39.50
1653protocol for access control and charging, but these are
1654not currently implemented in WAIS.
1655
1656The WAIS extensions to Z39.50 are mainly to provide the
1657relevance feedback capability.
1658
1659Note that the Z39.50 protocol is not stateless - the
1660result set may in some circumstances be retained by the
1661server for the user to further refine or refer to.
1662However, the subset of Z39.50 used by current WAIS
1663implementations mean that server implementations may be
1664stateless.
1665
1666Document type is determined by the server from
1667information in the database index (see below), and is
1668sent to the client as part of the result set.
1669
1670 WAIS Publishing
1671
1672The first step in preparing data for publishing in a WAIS
1673database is to use the 'waisindex' utility. This takes a
1674set of text files, and produces an index file which
1675contains an occurrence list of words of three or more
1676letters in every file. This index file is used by the
1677WAIS server software to resolve search requests from
1678clients.
1679
1680The 'waisindex' utility indexes files in a wide range of
1681text formats, as well as postscript and image files in
1682various encodings (only the file name is indexed for
1683image files). Some of the text formats involve a file as
1684being treated as a collection of documents for the
1685purposes of WAIS access. Note that there appears to be
1686no formal "registry of types" - just whatever the
1687waisindex program supports. There is no distinction
1688between media type and encoding format.
1689
1690 Published Non-text Data
1691
1692There is relatively little non-text data available in
1693WAIS databases.
1694
1695 URL=wais://quake.think.com:210/CM-images is a database
1696 of TIFF images from the Connection Machine.
1697
1698 URL=wais://mpcc3.rpms.ac.uk:210/home/images/pathology/
1699 RPMS-pathology is a database of histo-pathological
1700 images and documentation on mammalian endocrine
1701 tissue.
1702
1703 URL=wais://starhawk.jpl.nasa.gov:210/pio contains GIF
1704 images from NASA planetary probe missions, together
1705 with their captions. The presence of the caption
1706 index information makes it difficult to construct a
1707 search which returns images in the result set -
1708 increasing the maximum result set size may help.
1709
1710 Advantages
1711
1712WAIS is ideally suited for its intended purpose of
1713searching databases of textual information on the basis
1714of keywords. It appears to have the potential to satisfy
1715the requirements of some of the "database" category of
1716applications mentioned in Chapter 1.
1717
1718 Limitations
1719
1720WAIS is not (and does not pretend to be) a general-
1721purpose information system, as Gopher and WWW are. WAIS
1722does not have hyperlinking, and offers a purely flat
1723structure.
1724
1725A limitation which is particularly apparent is the way
1726that the current version of FreeWAIS indexes non-text
1727files - using only the filename! However, it does seem
1728that simply changing the indexing program to allow a list
1729of keywords to be attached to non-text files would
1730suffice to allow sensible indexing of non-text data. The
1731commercial (WAIS Inc) version of WAIS allows several
1732files to be associated together for indexing and
1733retrieval purposes. Furthermode, the UCSF Centre for
1734Knowlege Management is modifying the FreeWAIS code to
1735support the indexing of multiple content types. The
1736document returned by WAIS will be an HTML document
1737containing pointers to the multimedia data. Contact
1738dcmartin@library.ucsf.edu for further information.
1739
1740WAIS is not a fully-featured query/response protocol such
1741as SQL. It has no concept of fields, or numeric data
1742types.
1743
1744It appears to be impossible to retrieve a document from
1745its catalogue entry in many of the existing databases.
1746
1747 World-Wide Web
1748
1749The World-Wide Web project (also known as WWW or W3),
1750started and driven by CERN, is a large-scale distributed
1751hypertext system. It uses the standard client-server
1752paradigm, with client "browser" software responsible for
1753fetching and displaying data. Originally aimed at the
1754High Energy Physics community, it has spread to other
1755areas.
1756
1757Browser software is available for a large number of
1758systems including:
1759
1760 Line-mode dumb terminal.
1761 Terminal with Curses support
1762 Macintosh
1763 X/Motif
1764 X11
1765 PC/MS Windows
1766 NeXT
1767
1768There is server software available for:
1769
1770 VM mainframes.
1771 UNIX
1772 Macintosh
1773 VMS
1774
1775 WWW User Image
1776
1777The WWW world consists of nodes (usually called
1778documents) and links. Links are connections between
1779documents: to follow a link, a reader clicks with a mouse
1780on a word in the source document, which causes the linked-
1781to document to be retrieved and displayed. (On systems
1782without a mouse, the user types a number instead.)
1783
1784Indexes are special documents which, rather than being
1785read, may be searched. To search an index, a reader
1786supplies keywords (or other search criteria). The result
1787of a search is a "virtual" document containing links to
1788the documents found. All documents, whether real,
1789virtual or indexes, look similar to the reader.
1790
1791The WWW addressing mechanism means that an interface to
1792Gopher and anonymous FTP information sources may be
1793established, in a way which is transparent to the user.
1794Thus, the whole of gopherspace is part of the Web.
1795Transparent gateways to other systems, including Hyper-G
1796and WAIS, are also available.
1797
1798 URL
1799
1800All nodes on the Web are addressed using the "Universal
1801[or Uniform] Resource Locator" (URL) syntax, defined in
1802[Lee93a]. This is an Internet Draft produced by the IETF
1803URL Working Group.
1804
1805A URL is a name for an object (which may be a document or
1806an index) on the Internet. It has the general form:
1807
1808 <scheme> : <path> [ # <anchorid> ]
1809
1810The <scheme> identifies an access protocol or method for
1811the object. Some of the schemes are HTTP (the native WWW
1812protocol), anonymous FTP, Andrew file system, news, WAIS,
1813Gopher. The <path> component locates the document in a
1814way significant for the access method. Thus for instance
1815for anonymous FTP, the path includes the fully-qualified
1816domain name of the host on which the document resides,
1817and the directory and file name under which it may be
1818found. For some schemes, the <path> may include a search
1819string (or combination of strings) which is used to
1820address a "virtual" object formed by searching an index
1821of some kind. The HTTP, WAIS and Gopher schemes can use
1822search strings, which usually follow the rest of the
1823path, separated from it by a ?.
1824
1825The optional <anchorid> is used for addressing within an
1826object. Its interpretation is not defined in the URL
1827specification.
1828
1829"Partial" URLs may be specified. These are used within a
1830document on the Web to refer to another "nearby" document
1831- for instance to a document in another file on the same
1832machine. Certain parts of the URL (eg the scheme and
1833machine name) may be omitted, according to well-defined
1834rules. This makes it much easier to move groups of
1835documents around, while maintaining the links within and
1836between them.
1837
1838A URL locates one and only one object on the Internet.
1839However, more than one URL may point to the same object.
1840Given two URLs, it is not in general possible to
1841determine whether they refer to the same object.
1842Furthermore, there is no guarantee that a single URL will
1843refer to the same object at different times (the object
1844may change incrementally, or it may be completely
1845replaced with something different, or it may indeed be
1846removed).
1847
1848 HTTP
1849
1850HTTP (HyperText Transfer Protocol) is the protocol
1851employed between server and client. It is defined in
1852[Lee92]. The protocol is currently being revised (see
1853the Future Developments section below), and will
1854eventually be proposed as an Internet standard.
1855
1856The original protocol is extremely simple, and requires
1857only a reliable connection-oriented transport service,
1858typically TCP/IP. The client establishes a connection
1859with the server, and sends a request containing the word
1860GET, a space, and the partial URL of the node to be
1861retrieved, terminated by CR LF. The server responds with
1862the node contents, comprising a text document in the
1863Hypertext Markup Language (HTML). The end of the
1864contents is indicated by the server closing the
1865connection.
1866
1867 HTML
1868
1869HTML (HyperText Markup Language) is the way in which text
1870documents must be structured if they are to contain links
1871to other documents. Non-HTML text documents may of
1872course be made available on the Web, but they may not
1873contain links to other documents (ie they are leaf
1874nodes), and they will be displayed by browsers without
1875formatting, probably using a fixed-width font. Like
1876HTTP, HTML is also undergoing enhancement, but the
1877original version is defined in [Lee92], and is being
1878submitted as an Internet draft.
1879
1880HTML is an application of SGML (Standard Generalized
1881Markup Language). It defines a range of useful tags for
1882indicating a node title, paragraph boundaries, headings
1883of several different levels, highlighting, lists, etc.
1884Anchors are represented using an <A> tag. For instance,
1885here is an example of HTML containing an anchor:
1886
1887The HTTP protocol implements the WWW <A NAME=13
1888HREF="../../Administration/DataModel.html">data model</A>
1889.
1890
1891The location of the anchor is the text "data model". It
1892is a source anchor, with a target given by the URL in the
1893HREF attribute, so the text would appear highlighted in
1894some way in a client's window, to indicate that clicking
1895on it would cause a hyperlink to be traversed. It is
1896also a target anchor, with an anchor ID given by the NAME
1897attribute. A source anchor referring to this target
1898would specify #13 at the end of the node's URL.
1899Traversing a hyperlink to this node would cause the
1900entire node to be retrieved, but the target anchor text
1901would be displayed in some emphasised way - for instance
1902if the retrieved text is displayed in a scrolling window,
1903it might be positioned such that the target anchor
1904appears at the top of the window.
1905
1906Another attribute of the <A> element, TYPE, is also
1907available, which is intended to describe the nature of
1908the relationship modelled by the link. However, this is
1909not in extensive use, and there appears to be no registry
1910of the possible values of such types.
1911
1912 Future Developments
1913
1914HTTP and HTML are currently being extended in a backward-
1915compatible way to add multimedia facilities. [Lee93b]
1916describes the HTTP2 protocol. The revised HTML is
1917defined in [Lee93c]. Both documents are subject to
1918change (and indeed the HTML2 specification has changed
1919substantially during the preparation of this report).
1920
1921The revised HTML contains many enhancements which are
1922useful for multimedia support. Some of the most relevant
1923are listed below.
1924
1925 "Universal Resource Numbers" are a proposed system for
1926 unique, timeless identifiers of network-accessible
1927 files presently being designed by IETF Working Groups.
1928 URNs must be distinguished from URLs, which contain
1929 information sufficient to locate the document. URNs
1930 may be allocated to nodes and may be represented in
1931 source anchors. This saves client software from
1932 retrieving a copy of something it already has -
1933 allowing sensible caching of large video clips, for
1934 instance. The disadvantage is that when something is
1935 changed and given a new URN, the source anchors of all
1936 links which point to it must be changed (and the URNs
1937 of these documents must therefore be changed, and so
1938 on). Therefore, it makes sense to allocate URNs only
1939 to very large documents which change rarely, and not to
1940 the documents which reference them.
1941
1942 The title of a destination document may be included as
1943 an attribute of a source anchor. This allows a client
1944 to display the title to the user before or during
1945 retrieval, and also allows data which does not itself
1946 contain a title (eg image data) to be given one.
1947
1948 There is provision for in-line non-text data (eg
1949 images, video, graphics, mathematical equations), which
1950 appears in the same window as the main textual material
1951 in the node.
1952
1953 The concept of the relationship expressed by a
1954 hyperlink is expanded. Both source and target anchors
1955 may contain relation attributes which point forwards
1956 and backwards respectively. Possible relationships
1957 include "is an index for", "is a glossary for",
1958 "annotates", "is a reply to", "is embedded in", "is
1959 presented with". The last two are useful for
1960 multimedia - for instance, the "embed" relationship
1961 could cause a retrieved image to be fetched and
1962 embedded in the display of a text node, and the
1963 "present" relationship could cause a sound clip to be
1964 automatically retrieved and presented along with a text
1965 node.
1966
1967The HTTP2 protocol maintains the same stateless
1968connect/request/response/close procedure as the current
1969HTTP protocol. Data is transferred in MIME-shaped
1970messages, allowing all MIME data formats (including HTML)
1971to be used. As well as the GET operation, HTTP2 has
1972operations such as:
1973
1974HEAD Fetch attribute information about a node
1975 (including the media type and encoding)
1976
1977CHECKOUT/CHECKIN/PUT/POST
1978 These allow nodes to be checked out for
1979 updating and checked back in again, and
1980 new nodes to be created. New node data is
1981 supplied in MIME shape with the request.
1982
1983The request from the client can contain a list of formats
1984which the client is prepared to accept, user
1985identification, authorisation information (a placeholder
1986at present), an account name to charge any costs to, and
1987identification of the source anchor of the hyperlink
1988through which the node was accessed.
1989
1990The response from the server may contain a range of
1991useful attributes (eg date, cost, length - but only for
1992non-text data). The server may redirect the query,
1993indicating a new URL to use instead. It may also refuse
1994the request because of authorisation failure or absence
1995of a charge account in the request.
1996
1997The protocol also contains a mechanism which is designed
1998to allow the server to make an intelligent decision about
1999the most appropriate format in which to return data,
2000based on information supplied in the request by the
2001client. This may for instance allow a powerful server to
2002store the uncompressed bitmap of an image, but to
2003compress it on request using an appropriate encoding,
2004according to the decoding capabilities announced by the
2005client.
2006
2007An HTTP2 server and client are currently under test.
2008Some HTML2 features are already fitted to the XMosaic
2009browser.
2010
2011 Mosaic
2012
2013The Mosaic project, located at the US National Centre for
2014Supercomputing Applications (NCSA) at the University of
2015Illinois, is developing a networked information system
2016intended for wide-area distributed asynchronous
2017collaboration and hypermedia-based information discovery
2018and retrieval. Mosaic, which is specifically oriented
2019towards scientific research workers, has adopted the
2020World-Wide Web as the core of the system, and the first
2021Mosaic software to appear was the XMosaic WWW client for
2022UNIX with X. Other clients of similar functionality are
2023under development for the Apple Macintosh and the PC with
2024Windows.
2025
2026The capabilities of the XMosaic browser include:
2027
2028 Support for NCSA's Data Management Facility (DMF) for
2029 scientific data.
2030
2031 Support for transferring data with other NCSA tools
2032 such as Collage, using NCSA's Data Transfer Mechanism
2033 (DTM).
2034
2035 The ability to "check out" documents for revision, and
2036 to check them back in again.
2037
2038 Local and remote annotation of Web documents.
2039
2040Future planned functionality includes:
2041
2042 In-line non-text data (in addition to images).
2043
2044 Information space graphical representation and control.
2045
2046 Hypermedia document editing.
2047
2048 Information filtering.
2049
2050NCSA intends to make the entire Mosaic system publicly
2051available and distributable.
2052
2053The XMosaic browser was used extensively for finding and
2054retrieving information used to prepare this report.
2055
2056 Web Publishing
2057
2058Making a web is as simple as writing a few SGML files
2059which point to your existing data. Making it public
2060involves running the FTP or HTTP daemon, and making at
2061least one link into your web from another. In fact, any
2062file available by anonymous FTP can be immediately linked
2063into a web. The very small start-up effort is designed to
2064allow small contributions.
2065
2066At the other end of the scale, large information
2067providers may provide an HTTP server with full text or
2068keyword indexing. This may allow access to a large
2069existing database without changing the way that database
2070is managed. Such gateways have already been made into
2071Digital's VMS/Help, Technical University of Graz's "Hyper-
2072G", and Thinking Machine's WAIS systems.
2073
2074There are a few editors which understand HTML - for
2075instance on UNIX and on the NeXT platform.
2076
2077 Published non-text data
2078
2079See the multimedia demo node on:
2080
2081http://hoohoo.ncsa.uiuc.edu:80/mosaic-
2082docs/multimedia.html
2083
2084This contains links to images, sound, movies and
2085postscript media types. The media type is determined by
2086the filename extension in the URL specification of the
2087target node. The (XMosaic) client uses this to invoke a
2088separate program appropriate for displaying the media
2089type, or in some cases it can be displayed embedded
2090within the source document. The latter method uses an
2091<IMG> tag, which is part of HTML2.
2092
2093 Advantages
2094
2095WWW is a hypertext system and its underlying technology
2096is thus richer than Gopher. The use of SGML, which is of
2097increasing importance in hypermedia systems, allows a
2098great deal of expressiveness and structure, and enables
2099text to be presented in an attractive way. The
2100facilities for multimedia data in the extended versions
2101of HTTP and HTML are excellent. It also seems that QOS
2102and management issues identified in Chapter 2 are to some
2103degree catered for in these extensions.
2104
2105 Limitations
2106
2107There is no indication in the source anchor of the media
2108type of the destination node, or of its size (this has
2109been ruled out on the argument that the information is
2110likely to degrade with time). It is necessary to perform
2111a HEAD request (in HTTP2) to deduce this.
2112
2113Link source anchors must be in text documents, so non-
2114text nodes must be leaf nodes. However, with HTML2 using
2115the <IMG> tag, an embedded bitmap may be used as a source
2116anchor, and the position of the mouse click within the
2117image is passed to the server, which can then choose to
2118return a different document depending on where in the
2119image the mouse was clicked.
2120
2121WWW is much less prevalent than Gopher, partly because of
2122an (erroneous?) perception that setting up an HTTP server
2123is more complex than setting up a Gopher server. There
2124are only about 60 servers world-wide; however the growth
2125in the use of WWW is much faster than the growth in the
2126use of Gopher. The availability of sophisticated WWW
2127clients such as XMosaic is fuelling this growth.
2128
2129 Evaluating Existing Tools
2130
2131This section compares the capabilities of the Gopher,
2132WAIS and World-Wide Web systems (abbreviated as GWW) to
2133the informal requirements defined in section 2.3.
2134
2135 Platforms
2136
2137The table below gives the names of the most important
2138client software for each of GWW on the three most
2139important platforms of interest. WWW is the weakest,
2140with clients for the Macintosh and the PC still under
2141development. The main PC Gopher client is "PC Gopher
2142III", which is a DOS program, not a Windows program.
2143
2144
2145
2146CLIENTS Gopher WAIS WWW
2147
2148Macintosh TurboGopher WAIStation (No name)
2149 (beta
2150 version
2151 available)
2152
2153PC with HGopher WAIS for Cello (beta
2154Windows (two others Windows, version
2155 also WAIS available),
2156 available) Manager Mosaic
2157 (beta due
2158 3Q93)
2159
2160UNIX with X Xgopher, XWAIS XMosaic
2161 XMosaic
2162
2163
2164
2165At present, multimedia support in most of these clients
2166(where it exists) is limited to the invocation of
2167external "viewer" programs for particular media types.
2168The exception is XMosaic, which supports in-line images
2169in WWW documents.
2170
2171 Media Types
2172
2173The GWW tools can all handle multiple media types well.
2174
2175 Text is very well supported by all three tools. WWW
2176 offers facilities for displaying "richer" text,
2177 supporting headings, lists, emphasised text etc, in a
2178 standardised way.
2179
2180 Image data is also well supported, using either
2181 external viewers (eg the TurboGopher client software on
2182 a Macintosh might invoke the JPEGView program to
2183 display an image); or in-line display within a text
2184 document (WWW with XMosaic on UNIX).
2185
2186 There is little direct support for application-specific
2187 data, but most systems allow data of a nominated type
2188 to be passed to an external viewer or editor program.
2189 This tends to be a function of the client software
2190 rather than being built in to the protocol or server.
2191 There has been discussion in the WWW community about
2192 using TeX for representing mathematical equations, and
2193 about providing "panels" within a text document where a
2194 separate application could render its application-
2195 specific data (or indeed any data which can be
2196 represented spatially). This latter suggestion fits
2197 well with the OLE (Object Linking and Embedding)
2198 approach used in Microsoft Windows.
2199
2200 Sound can be supported through the external "viewer"
2201 concept. Some platforms don't have readily-available
2202 "viewers" with "tape recorder"-style controls for
2203 replaying. There is no single commonly-accepted sound
2204 encoding format.
2205
2206 Video data can be handled using external viewers. MPEG
2207 and QuickTime are the most common encodings.
2208
2209One essential capability of a client/server protocol is
2210the ability for the client to determine the type of a
2211node (and a list of available encodings) before
2212downloading it. WAIS and Gopher transfer this
2213information in the result set and menu respectively. WWW
2214clients currently determine this information either from
2215analysing the URL of a target node, or by the occurrence
2216of the <IMG> tag. The new WWW HTTP2 protocol allows the
2217media type and encoding of a node to be determined
2218through a separate interaction with the server.
2219
2220The GWW systems all use different methods for expressing
2221type and encoding. WAIS does not distinguish the
2222encoding from the media type. WWW is moving to the MIME
2223type/encoding system. Gopher does not distinguish type
2224and encoding, but Gopher+ does, and is also moving to the
2225MIME type/encoding system.
2226
2227 Hyperlinks
2228
2229Only the WWW system has hyperlinks. Source anchors may
2230be text, images, or points within an image. Target
2231anchors may be entire nodes of any media type, or points
2232within (with HTTP2, portions of) text nodes.
2233
2234Gopher+ could potentially be enhanced to include
2235hyperlinks, but there seems to be no development effort
2236going towards this - those who need hyperlinking are
2237using WWW.
2238
2239Gopher menus can be constructed to allow alternative
2240views of gopherspace. For instance, a geographically-
2241organised menu tree of gopherspace is in place, but a
2242parallel subject-based menu tree could be added as an
2243alternative way of access to the same data. (There are
2244in fact moves to set this up.) Since WWW offers a
2245superset of Gopher functionality, these comments also
2246apply to the Web. In fact, the Web already has a
2247rudimentary subject tree.
2248
2249In both Gopher and WWW, non-textual data may be used in
2250different information structures without having to
2251maintain more than one copy.
2252
2253 Presentation
2254
2255There is little support in GWW for controlling the
2256presentation of non-text data.
2257
2258 Backdrops are not supported by GWW.
2259
2260 Buttons are supported in a limited way - typically, a
2261 node is retrieved by clicking on a highlighted text
2262 phrase, or on an entry in a list. In XMosaic, bitmap
2263 images can be used as buttons. However, there is no
2264 support for different styles of button. Client
2265 software may have generic navigation buttons (eg
2266 "Back", "Next", "Home") which are always available and
2267 don't form part of a node.
2268
2269 Synchronisation in space is not supported by GWW,
2270 except that WWW supports contextual synchronisation of
2271 images using the <IMG> tag.
2272
2273 Synchronisation in time is not supported by GWW.
2274
2275 Searching
2276
2277WAIS supports keyword searching, and is very well suited
2278for that task. The Gopher+ protocol could potentially
2279support multimedia database querying applications through
2280the ASK attribute, but there is as yet no server
2281implementation which supports such database applications.
2282In the WWW project, there are ongoing discussions on how
2283best to extend HTML to cope with database query
2284applications - an <INPUT> tag has been suggested - but no
2285consensus has yet emerged.
2286
2287Both Gopher and WWW can make use of WAIS-type keyword
2288searching: either by incorporating WAIS code into the
2289server (enabling WAIS index files to be searched); or
2290through WAIS gateways, which run searches on remote WAIS
2291servers in response to queries from non-WAIS clients.
2292
2293 Interaction
2294
2295XMosaic allows users to make text (or on some platforms,
2296audio) annotations to any text node. The annotations
2297appear at the end of the text display. They are held
2298locally - other users of the node do not see the
2299annotations (but a recently added facility allows
2300globally-visible annotations held on an "annotation
2301server"). Text annotations may include hyperlinks to
2302other nodes (provided the user knows how to use HTML).
2303Other clients do not provide such facilities.
2304
2305There is a move to add an "email" address notation to
2306URL. This would allow WWW client software to invoke a
2307mail program when a user selects an anchor with such a
2308URL.
2309
2310There are plans to allow WWW users to delineate a
2311rectangular area of interest within an image for use in
2312an HTTP request.
2313
2314There is no support in GWW clients for interacting with
2315sequences of images in the way described in section
23162.3.6.
2317
2318 Quality of Service
2319
2320The user expectations for responsiveness mentioned in
2321section 2.3.7 are difficult to meet with currently-
2322deployed wide-area network (or even LAN) technology,
2323particularly for voluminous multimedia data. None of the
2324GWW systems currently exploit the emerging isochronous
2325data transfer capabilities of protocols such as RTP and
2326technologies such as ATM. None of them make serious
2327attempts to alleviate the problem in other ways (except
2328for WWW, which defines some mechanisms in HTTP2 for
2329format negotiation based on size and available bandwidth
2330considerations).
2331
2332 Management
2333
2334The following table shows the support for three key
2335management facilities in the GWW systems. The first two
2336facilities require support in the client/server protocol,
2337the third requires support in the server, but depends on
2338authentication being available.
2339
2340 Gopher WAIS WWW
2341
2342Access No No1 Yes, in
2343control and HTTP2
2344authenticatio
2345n
2346
2347Charging No No Yes, in
2348support HTTP2
2349
2350Monitoring No No No
2351for
2352statistical
2353and
2354assessment
2355purposes
2356
2357
2358
2359Note:
2360
23611."Access-control-facility" is a feature of Z39.50 which
2362 is not used by the current WAIS implementations.
2363
2364 Scripting Requirements
2365
2366None of the GWW systems have facilities for the execution
2367of scripts by the client, because of security issues (it
2368would be too easy for a malicious "trojan" script to be
2369executed). Gopher and WWW servers have the ability for a
2370UNIX script to be run by the server, with the script
2371output returned to the client. Scripting as understood
2372in the context of stand-alone multimedia applications
2373does not exist in GWW.
2374
2375 Bytestream Format
2376
2377None of the three GWW systems use a bytestream format for
2378interchanging collections of material. There has been
2379some talk about setting up a system akin to the "Trickle"
2380mail server, for retrieving single document nodes from
2381GWW using mail. Such a system has been implemented for
2382WWW.
2383
2384 Authoring tools
2385
2386Gopher is sufficiently simple to set up that no special
2387authoring tools are required. WAIS requires only an
2388indexing program (as discussed in section 3.2) for
2389preparing material for publication.
2390
2391WWW, because it uses a sophisticated authoring language
2392(HTML), benefits from the availability of authoring
2393tools. There are HTML editors for UNIX (using the tk
2394toolkit) and the NeXT system. There are no authoring
2395tools designed specifically for exploiting the multimedia
2396capabilities of WWW, mainly because these capabilities
2397are still evolving.
2398
2399 Research
2400
2401This section describes some current research projects in
2402the area of distributed hypermedia information systems.
2403
2404 Hyper-G
2405
2406Hyper-G [KS92] is an ambitious distributed hypermedia
2407research project at a number of institutes of the IIG
2408(Institutes for Information-Processing Graz), the
2409Computing and Information Services Centre of the Graz
2410University of Technology, and the Austrian Computer
2411Society. It is funded by the Austrian Ministry of
2412Science. It combines concepts of hypermedia, information
2413retrieval systems and documentation systems with aspects
2414of communication and collaboration, and computer-
2415supported teaching and learning.
2416
2417Unlike WWW, Hyper-G supports bi-directional links. This
2418enables users to see which other documents reference the
2419one they are using, and also allows the system to avoid
2420dangling pointers when a linked-to document is deleted.
2421Another difference from WWW is that links are kept
2422separately from their source and target nodes, to allow
2423easy linking of read-only documents and for ease of link
2424maintenance. In addition to manually defined links,
2425Hyper-G supports automatic static and dynamic (ie view-
2426time) generation and maintenance of links.
2427
2428Hyper-G has a concept of generic "structures" - an
2429additional layer of relationships imposed on (and
2430orthogonal to) the web of documents and links. A
2431document can be part of more than one structure, and
2432structures may be hierarchically related. Types of
2433structure include:
2434
2435 "Clusters" are a set of documents which are all
2436 presented together.
2437
2438 "Collections" are unordered sets of documents or other
2439 structures, and can be used as query domains or to
2440 construct gopher-like menus.
2441
2442 "Paths" are ordered sets of documents or structures,
2443 which must be visited sequentially
2444
2445One application of the structure concept is the provision
2446of "guided tours" through the information space.
2447
2448In addition to hypernavigation, the collection hierarchy
2449and guided tours, another strategy for interaction with
2450the system is the use of database queries. Two kinds of
2451query are supported: keyword searching in a user-defined
2452list of databases; and collection-specific form-filling
2453queries. In the latter case, the answer to the query may
2454appear dynamically as the form is filled out.
2455
2456Four modes of user identification are supported:
2457"identified", where a userid is publicly associated
2458through name and address information with a particular
2459individual; "semi-identified", where a userid is
2460associated by the system with an individual, but the user
2461is only known to other users through a pseudonym;
2462"anonymously identified", where the userid is not
2463associated by the system with any individual; and
2464"anonymous", where there is no userid (or a generic
2465userid such as "guest"). Possible operations in the
2466system depend on the user's mode of identification.
2467Users may access the system in any desired mode, and
2468switch to other modes only when necessary.
2469
2470Hyper-G contains specific support for multilingual
2471documents and document clusters. Users may specify an
2472ordered list of preferred languages, for instance. There
2473are plans to experiment with automatic translation
2474programs.
2475
2476Integration of other, external, systems such as WWW into
2477Hyper-G in a seamless manner is possible.
2478
2479Hyper-G is in use as a CWIS within Graz Technical
2480University. Client software is available for UNIX
2481workstations from DEC, HP, SGI, and SUN. The system is
2482still in an experimental state, but it has been used by
2483about 200 students as part of a course on the social
2484impact of information technology.
2485
2486 Microcosm
2487
2488Microcosm [DHH92] is an open hypermedia system developed
2489at the University of Southampton. It is implemented on
2490the PC under MS Windows, and versions for the Apple
2491Macintosh and for UNIX with X are under development.
2492
2493Microcosm consists of a number of autonomous processes
2494which communicate with each other by a message-passing
2495system. Information about hyperlinks between documents
2496is stored in a link database, or "linkbase", and is not
2497stored in the documents themselves. This has the
2498advantages that:
2499
2500 Links to and from read-only documents (perhaps stored
2501 on CD-ROM) are possible.
2502
2503 Documents need undergo no conversion process to be
2504 imported into the system - they can still be viewed and
2505 edited using the original application which created
2506 them, without the link information getting in the way.
2507
2508 It is as easy to establish links to and from non-text
2509 documents as text documents.
2510
2511In Microcosm, the user interacts with a "viewer" program
2512for a particular media type. Such programs may be
2513specifically written for use with Microcosm (about 10
2514such viewers have been written for a number of common
2515media types and encodings); or they may be a program
2516adapted for use with Microcosm (the programmability of
2517Microsoft Word for Windows has allowed it to be so
2518adapted); or it may even be a program with no knowledge
2519of Microcosm.
2520
2521The user selects an object (eg a piece of text) in the
2522viewer, and requests Microcosm to perform an action with
2523the object - typically to follow a link to another
2524document. This may involve executing another viewer to
2525display the target document.
2526
2527Microcosm link source anchors may be specific (denoting a
2528unique point in a particular document), local (denoting
2529any occurrence of a particular object in a particular
2530document) or generic (denoting any occurrence of an
2531object in any document). Target anchors may specify
2532specific objects within a document. Other link styles
2533are text-retrieval links (looking up a full-text index ,
2534as WAIS does), and relevance links to a set of documents
2535using similar vocabulary to the source document (again,
2536similar to WAIS's relevance feedback).
2537
2538Links may be created by readers as well as by authors.
2539Dynamically-computed links may be added to the permanent
2540linkbase for later use. A history of link traversal is
2541maintained, and "guided tours" may be established through
2542the system which allow the reader to stray from and
2543return to the tour.
2544
2545Microcosm viewers operate by sending messages to the
2546Microcosm system. In MS Windows, these messages are
2547transferred using DDE (Dynamic Data Exchange); in the
2548Apple Macintosh version Apple Events are used, and
2549sockets are used on UNIX. For viewers which are not
2550Microcosm aware, the user must transfer the selected
2551object to the system clipboard before being able to
2552follow a link from it.
2553
2554Networking support in Microcosm is currently under
2555development. Components of Microcosm may be distributed
2556to multiple machines - there is not necessarily a concept
2557of "client" and "server".
2558
2559There are problems with the Microcosm approach, common to
2560systems which maintain link information separately from
2561documents, and which use external viewers.
2562
2563 Documents move and change, thus invalidating links.
2564 Microcosm datestamps links to help to detect (but not
2565 correct) such problems.
2566
2567 It is not always clear what links are available to be
2568 followed from a document, since the viewer program is
2569 unaware of the contents of the linkbase.
2570
2571 It is not always possible to indicate the object within
2572 a document which is the target anchor of a link. Many
2573 viewers automatically show the start of the document
2574 (eg a word processor), or perhaps the entire document
2575 (eg a picture viewer). The user has no way of knowing
2576 which part of the target document the link just
2577 followed points to.
2578
2579Microcosm may be viewed as an integrating hypermedia
2580framework - a layer on top of a range of existing
2581applications which enables relationships between
2582different documents to be established.
2583
2584Microcosm is currently being "commercialised".
2585
2586 AthenaMuse 2
2587
2588AthenaMuse 2 (AM2) is an ambitious distributed hypermedia
2589authoring and presentation system under development by
2590the AthenaMuse Software Consortium based at MIT. It is
2591based on the earlier AM1 system developed as part of
2592MIT's Project Athena. The first version of AM2 is
2593scheduled for January 1994, and will be "pre-commercial
2594software", with a fully-commercialised version due about
25956 months later. Both the educational and commercial
2596sectors are the intended market. The system will
2597initially be based on X and UNIX workstations, but
2598PC/Windows will also be supported in a second phase.
2599Apple Macintosh support has a lower priority.
2600
2601The specifications of AM2 are available in [BCH92]. Some
2602of the key points are:
2603
2604 AM2 will support import and export of application from
2605 and to standard forms. The project is watching
2606 standards such as HyTime, MHEG and ODA.
2607
2608 Several "application themes", or frequently-occurring
2609 collections of functionality, are viewed as useful.
2610 These are as follows:
2611
2612Application Theme Interac
2613 tive?
2614Presentation of multimedia data No
2615Exploration of a rich Yes
2616multimedia environment
2617Simulation of a real-world Partial
2618scenario ly
2619Communication of real-time No
2620information to the user
2621Authoring Yes
2622Annotation of material Yes
2623
2624
2625 "Interface templates" allow a multimedia application to
2626 make use of a common format for presenting a range of
2627 content. This is similar to the "backdrop" concept
2628 mentioned in section 2.3.4.
2629
2630 A range of link types will be supported.
2631
2632 Media content editors and interface/application editors
2633 for structuring will be provided. A third class of
2634 editor, the "hypermedia notebook", will allow readers
2635 to excerpt and annotate media from AM2 applications.
2636
2637The project is developing multimedia network services,
2638including the transmission of digital video, using a
2639client-server paradigm.
2640
2641 CEC Research Programmes
2642
2643Some of the research programmes sponsored by the
2644Commission for the European Community (CEC) contain
2645apparently relevant projects. [Adi93] has further
2646details of some of these projects.
2647
2648RACE programme
2649
2650The RACE programme is outlined in [CEC92a], which should
2651be consulted for further information about the projects
2652described below. The RACE programme targets the
2653industrial, commercial and domestic sectors, and results
2654are not necessarily directly applicable to the research
2655and academic community. RACE project numbers are given.
2656
2657RACE Phase I projects, which have mostly completed:
2658
2659R1038 MCPR - Multimedia Communication, Processing and
2660 Representation
2661 This project developed a demonstrator multimedia
2662 system with communications capability for travel
2663 agents.
2664
2665R1061 DIMPE - Distributed Integrated Multimedia
2666 Publishing Environment
2667 The project designed and implemented interim
2668 services for compound document handling, and defined
2669 a distributed publishing architecture.
2670
2671R1078 European Museums Network
2672 This project aimed to demonstrate interactive
2673 navigation through a pool of multimedia museum
2674 objects, using ISDN as the communications network.
2675
2676RACE Phase II projects:
2677
2678R2008 EuroBridge
2679 Aims to demonstrate multi-point multimedia
2680 applications running over DQDB, FDDI and ATM test
2681 networks.
2682
2683R2043 RAMA - Remote Access to Museum Archives
2684 This project follows on from R1078.
2685
2686R2060 CIO - Coordination, Implementation and
2687 Operation of Multimedia Services
2688 One aspect of this project is JVTOS - a "Joint
2689 Viewing and Teleoperation Service". This aims to
2690 integrate standard multimedia applications running
2691 on a range of heterogeneous machines into a
2692 cooperative working environment, allowing
2693 individuals to view and interact with multimedia
2694 data on colleague's machines.
2695
2696ESPRIT Programme
2697
2698The ESPRIT research programme is outlined in [CEC92b],
2699which should be consulted for further information about
2700the projects listed below. ESPRIT project numbers are
2701given.
2702
270328 MULTOS - A Multimedia Filing System
2704 This project, which ran from 1985 to 1990, developed
2705 a client/server system for filing and retrieval of
2706 multimedia documents using the ODA interchange
2707 format standard (ODIF).
2708
27095252 HYTEA - HyperText Authoring
2710 This project, which runs from 1991 to 1994, aims to
2711 develop a set of authoring tools for large and
2712 complex hypermedia applications.
2713
27145398 SHAPE - Second Generation Hypermedia Application
2715 Project
2716 This project is developing a portable software
2717 environment comparable to a CASE tool intended to
2718 facilitate the realisation of complex hypermedia
2719 applications.
2720
27215633 HYTECH - Hypertextual and Hypermedial Technical
2722 Documentation
2723 This project, which ran from 1990-1991, was to
2724 assess the feasibility of hypermedia technology and
2725 to devise needed extensions to it in order to
2726 support applications dealing with technical
2727 documentation management.
2728
27296586 PEGASUS - Distributed Multimedia Operating System
2730 for the 1990s
2731 This project is aimed at the design of an operating
2732 system architecture for scalable distributed
2733 multimedia systems and the development of a
2734 validating prototype, the design and implementation
2735 of a distributed complex-object service and a global
2736 name service, the development of mechanisms for the
2737 creation, communication and rendering of fully
2738 digital multimedia documents in real time and in a
2739 distributed fashion, and the design and
2740 implementation of an application for the system: a
2741 digital TV director.
2742
27436606 IDOMENEUS - Information and Data on Open Media for
2744 Networks of Users. This project, which started
2745 January 1993, brings together workers in the
2746 database, information retrieval, networking and
2747 hypermedia research communities in the development
2748 of an "ultimate information machine". It "will
2749 coordinate and improve European efforts in the
2750 development of next-generation information
2751 environments capable of maintaining and
2752 communicating a largely extended class of
2753 information on an open set of media". Because of
2754 the close match between the subject of the IDOMENEUS
2755 project and the RARE WG-IMM, it is recommended that
2756 RARE establish a liaison with this project.
2757
2758 Other
2759
2760Some other research projects of less immediate relevance
2761are listed below. Some of these projects are described
2762further in [Adi93].
2763
2764 Xanadu is a project to develop an "open, social
2765 hypermedia" distributed database server, incorporating
2766 CSCW features. It has been in existance for many years
2767 and has been funded by a number of companies. The
2768 current status of this project is not known, and
2769 although iminent availability of alpha-test versions
2770 has been announced more than once, no software has been
2771 delivered.
2772
2773 CMIFed [RJM93] is an editing and presentation
2774 environment for portable hypermedia documents being
2775 developed at CWI, Amsterdam, NL. It is based on the
2776 "Amsterdam Model" of hypermedia [HBR93], which is an
2777 extension of the Dexter hypertext reference model
2778 incorporating "channels" for media delivery and
2779 synchronisation constraints.
2780
2781 Deja Vu [Eli93] is a proposed "intelligent" distributed
2782 hypermedia application framework. It is intended as a
2783 vehicle for research in the areas of: hypermedia
2784 systems, object-oriented programming, distributed logic
2785 programming, and intelligent information systems.
2786 Proposed techniques for use in the Deja Vu framework
2787 include "inferential links", defined automatically
2788 according to predefined rules. A scripting language
2789 for use both by information providers and users is
2790 planned. This project is at a very early (proposal)
2791 stage, and as yet relatively little software has been
2792 developed. Deja Vu is intended principally as a
2793 research framework rather than as a service tool.
2794
2795 Demon is a project at Bellcore, US, investigating the
2796 network requirements of near-term residential
2797 multimedia services. The project is designing and
2798 implementing an experimental application which serves
2799 the needs of casual multimedia users.
2800
2801 InfoNote is a distributed, multiuser hypermedia system
2802 from Japan, implemented on a NEC EWS4800 running UNIX
2803 and X. InfoNote has an editor which can create
2804 Japanese texts, figures, and raster images. The same
2805 windows are used both for editors and browsers. The
2806 functionality of the window can be changed at any time
2807 if data is not write-protected.
2808
2809 MADE - Multimedia Application Demonstration Environment
2810 - is a project at British Telecom's research laboratory
2811 which centres on the use of the developing MHEG
2812 standard to access a multimedia object server. The
2813 server platform is a Sun SPARCstation with an object-
2814 oriented database package (ONTOS). Audio, video, text
2815 and graphical media types are covered. The University
2816 of Kent is working on a sub-project: "Multi-user
2817 Indexing in a Distributed Multimedia Database".
2818
2819 Zenith aimed to establish a set of principles to assist
2820 designers and developers of object management systems
2821 intended for distributed multimedia design
2822 environments. The project implemented a prototype
2823 generalised multimedia object management system.
2824
2825 Standards
2826
2827 Structuring Standards
2828
2829This section describes some of the important standards
2830for providing hyperstructure to multimedia data.
2831
2832 SGML
2833
2834SGML (Standard Generalized Markup Language - ISO 8879) is
2835a metalanguage for defining markup notations for text.
2836SGML is used to write Document Type Definitions or DTDs,
2837to which individual document instances must conform. It
2838finds application in a wide and increasing range of text
2839processing applications.
2840
2841The relevance of SGML to distributed hypermedia systems
2842is surprisingly high, mainly because of the great
2843expressive power of SGML, and its ability to handle non-
2844textual data using "external entities" and "notations".
2845
2846 The World-Wide Web is an SGML application with its own
2847 DTD.
2848
2849 The important HyTime hypermedia structuring standard
2850 (see below) is based on SGML.
2851
2852 The forthcoming MHEG hypermedia structuring standard
2853 (see below) has an SGML encoding.
2854
2855 SGML has been used in research hypermedia systems - for
2856 example Microcosm.
2857
2858 SGML is used in some commercial hypermedia systems -
2859 for example DynaText.
2860
2861 SGML is of increasing importance for academic
2862 publishing houses.
2863
2864It was interesting to note that at a recent (CEC-
2865sponsored) workshop on Hypertext and Hypermedia
2866standards, most of the speakers were conversant with and
2867supportive of the use of SGML for such systems.
2868
2869A related standard which may become important for SGML on
2870networks is SDIF (SGML Data Interchange Format - ISO
28719069). This standard specifies how an SGML document,
2872which may exist in a number of separate files of
2873different media types, may be encoded using ASN.1 into a
2874single bytestream. The entity structure is preserved, so
2875that the bytestream may be decoded by the recipient into
2876the same set of files.
2877
2878 HyTime
2879
2880HyTime (Hypermedia/Time-Based Structuring Language) is a
2881standardised infrastructure for the representation of
2882integrated, open hypermedia documents. It was developed
2883principally by ANSI committee X3V1.8M, and was
2884subsequently adopted by ISO and published as ISO 10744.
2885
2886HyTime is based on SGML. It is not itself an SGML DTD,
2887but provides constructs and guidelines ("architectural
2888forms") for making DTDs for describing Hypermedia
2889documents. For instance, the Standard Music Description
2890Language (SMDL: ISO/IEC Committee Draft 10743) defines a
2891(meta-)DTD which is an application of HyTime. In fact,
2892HyTime started as an attempt to produce a markup scheme
2893for music publishing purposes.
2894
2895HyTime specifies how certain concepts common to all
2896hypermedia documents can be represented using SGML.
2897These concepts include:
2898
2899 association of objects within documents with hyperlinks
2900
2901 placement and interrelation of objects in space and
2902 time
2903
2904 logical structure of the document
2905
2906 inclusion of non-textual data in the document
2907
2908An "object" in HyTime is part of a document, and is
2909unrestricted in form - it may be video, audio, text, a
2910program, graphics, etc. The terminology used in HyTime
2911(and in this section) thus differs slightly from the
2912terminology used in the rest of this report. A HyTime
2913object corresponds roughly to a node as defined in
2914section 1.2, and a HyTime document is a hyperdocument in
2915the terminology of this report.
2916
2917HyTime consists of six modules, which are very briefly
2918and selectively described below:
2919
2920 Base module. This provides facilities required by
2921 other modules, including a lexical model for describing
2922 element contents; facilities for identifying policies
2923 for coping with changes to a document, or traversing a
2924 link ("activity tracking"); and the ability to define
2925 "container entities" which can hold multiple data
2926 objects. This last was added to the HyTime standard at
2927 a late stage, at the instigation of Apple Computers
2928 Inc, as a "hook" for their Bento specification [HR92].
2929
2930 Measurement module. This allows for an object to be
2931 located in time and/or space (which HyTime treats
2932 equivalently), or any other domain which can be
2933 represented by a finite coordinate space, within a
2934 bounding box called an "event", defined by a set of
2935 coordinate points. Coordinates may be expressed in any
2936 units (predefined units include femtoseconds,
2937 fortnights, millenia, angstroms, Northern feet and
2938 lightyears!).
2939
2940 Location Address module. In addition to the
2941 fundamental ability of SGML to identify and refer to
2942 elements, this module provides a special "named
2943 location address" architectural form which can be used
2944 to refer indirectly to data which spans elements, or
2945 which is located in external entities. Data may also
2946 be addressed indirectly through the use of "queries",
2947 which return addresses of objects within some domain
2948 which have properties matching the query. A "HyQ"
2949 notation is provided for defining the query.
2950
2951 Hyperlinks module. Two basic types of hyperlink are
2952 defined: the contextual link (clink) has two anchors,
2953 one of which is embedded in a document to explicitly
2954 denote the anchor location; and the independent link
2955 (ilink) which may have more than two anchors, and which
2956 does not require the anchors to be embedded in the
2957 document. ilinks thus allow hyperlink information to
2958 be maintained separately from document content.
2959
2960 Scheduling module. This specifies how events in a
2961 source finite coordinate space (FCS) are to be mapped
2962 onto a target FCS. For instance, events on a time axis
2963 could be projected onto a spatial axis for graphical
2964 display purposes, or a "virtual" time axis as used in
2965 music could be projected onto a physical time axis.
2966
2967 Rendition module. This allows for individual objects
2968 to be modified before rendition, in an object-specific
2969 way. One example is modification of colours in image
2970 so that it can be displayed using the currently-
2971 selected colour map on a graphics terminal, or changing
2972 the volume of an audio channel according to a user's
2973 requirements.
2974
2975It is not envisaged that a hypermedia application would
2976need to use the entire range of HyTime facilities. An
2977application designer is able to choose appropriate HyTime
2978architectural forms, and to add application-specific
2979constraints to them. The designer may also of course use
2980non-HyTime SGML elements and attributes, but these
2981aspects of the application can't be understood by a
2982"HyTime engine". Even in the absence of a HyTime engine,
2983the HyTime architectural forms provide a useful base of
2984ideas from which a hypermedia system designer may wish to
2985work.
2986
2987The role of a HyTime engine is not specified in the
2988standard, but essentially it is a (sub)program which
2989recognises HyTime constructs in document instances and
2990performs application-independent processing on them. For
2991instance, it could interact with multimedia network
2992servers to resolve and access hyperlink anchors. A
2993commercial HyTime engine (HyMinder) is under development
2994by TechnoTeacher in the US, and the Interactive
2995Multimedia Group at the University of Massachusetts -
2996Lowell (contact lrutledg@cs.ulowell.edu) is also working
2997on a HyTime engine (HyOctane).
2998
2999The Davenport group (a loose consortium of interested
3000companies and individuals) is producing a series of
3001standards on hypermedia which further constrain the
3002HyTime architectural forms. One example is the SOFABED
3003module [NN93], which standardises the representation of
3004certain kinds of navigational information - tables of
3005contents, indexes and glossaries.
3006
3007HyTime was envisaged as an interchange format rather than
3008as a format for directly-executable hypermedia
3009applications. It is therefore very expressive, but may
3010be difficult to optimise for run-time efficiency.
3011
3012An attempt has been made [Kim93] to adapt the hyperlink
3013structure in WWW's existing HTML DTD to comply with
3014HyTime's clink architectural form. This requires changes
3015to WWW document instances as well as to browser software,
3016and in the absence of any immediate benefit it has found
3017little favour with the WWW community. However, it is
3018possible that HTML2 will use some aspects of HyTime.
3019
3020It is recommended that any further RARE work on networked
3021hypermedia should take account of the importance of SGML
3022and HyTime.
3023
3024 MHEG
3025
3026MHEG stands for the Multimedia and Hypermedia information
3027coding Experts Group, also known as ISO/IEC
3028JTC1/SC29/WG12 (it used to come under SC2). This group
3029is developing a standard "Coded Representation of
3030Multimedia and Hypermedia Information Objects" (ISO CD
303113522, or CCITT T.171), commonly called MHEG. The
3032standard is to be published in two parts - part 1 being
3033the base notation, representing objects using ASN.1, and
3034part 2 being an alternate notation which uses SGML. Part
30351 has nearly (June 1993) achieved CD status, and is
3036intended to reach full IS in 1994. Part 2 is intended to
3037reach the CD stage in late 1993.
3038
3039MHEG is suited to interactive hypermedia applications
3040such as on-line textbooks and encyclopaedia. It is also
3041suited for many of the interactive multimedia
3042applications currently available (in platform-specific
3043form) on CD-ROM. MHEG could for instance be used as the
3044data structuring standard for a future home entertainment
3045interactive multimedia appliance. Telecommunications
3046operators are interested in MHEG for providing
3047interactive multimedia services across ISDN.
3048
3049To address such markets, MHEG represents objects in a non-
3050revisable form, and is therefore unsuitable as an input
3051format for hypermedia authoring applications: its place
3052is perhaps more as an output format for such tools. MHEG
3053is thus not a multimedia document processing format -
3054instead it provides rules for the structure of multimedia
3055objects which permits the objects to be represented in a
3056convenient "final" form with the aim of direct
3057presentation.
3058
3059The MHEG draft standard is expressed in object-oriented
3060terms. The main object classes are outlined briefly
3061below.
3062
3063 Content class. A content object contains the encoded
3064 (monomedia) information to be presented, along with
3065 attributes which identify the type of information and
3066 the encoding method, and media-specific attributes such
3067 as fonts used, sampling rate, image size, etc.
3068
3069 Selection class and Modification class. The user may
3070 interact with MHEG objects which inherit interactive
3071 behaviour from these classes. (The MHEG object model
3072 supports multiple inheritance.)
3073
3074 Action class. Two types of action may be applied to
3075 objects: projection, which controls how objects are
3076 rendered; and status actions which affect the state of
3077 objects.
3078
3079 Link class. MHEG hyperlinks connect a "start" object
3080 with one or more "end" objects. Links consist of a set
3081 of conditions relating to the state of the start
3082 object, and a set of actions which are carried out when
3083 these conditions are satisfied. Links also define the
3084 spatio-temporal relationships between objects.
3085
3086 Script class. Script objects are used to describe more
3087 complex interobject linkages (eg multiple-source
3088 links). MHEG does not define a scripting language -
3089 instead it provides a formalism for encapsulating
3090 scripts which may be executed by an external program
3091 (see SMSL below).
3092
3093 Composite class. Related objects may be grouped
3094 together into a single composite object (recursively).
3095 The relationships between content objects within a
3096 composite object are determined by link and script
3097 objects which also are members of the composite object.
3098
3099 Descriptor class. Descriptor objects contain general
3100 information about sets of interchanged objects, so that
3101 a target system can ensure it has adequate resources to
3102 run the hypermedia application represented by the
3103 object set.
3104
3105The relationship between HyTime and MHEG has not yet been
3106fully established. One possible relationship [Mar91] is
3107that an MHEG application could be the output of a
3108compilation process which used an equivalent HyTime
3109document as input. This approach would benefit both from
3110the expressive power of HyTime and the run-time
3111efficiency of MHEG. However, it has yet to be shown that
3112this is feasible, since the capabilities of HyTime and
3113MHEG do not completely overlap.
3114
3115There seems to be relatively little interest in or
3116awareness of MHEG within the Internet community, which is
3117only just beginning to be aware of HyTime. In view of
3118the draft nature of the MHEG standard, this report
3119recommends that RARE should not invest substantial effort
3120in MHEG at this time. However, particularly in view of
3121the interest in it shown by PTTs, a watching brief should
3122be kept on MHEG, as it may well be relevant in the
3123future.
3124
3125 ODA
3126
3127The Open Document Architecture standard (ODA - ISO 8613
3128or T.140) is a compound document interchange format
3129designed for transferring documents between open systems.
3130It is able to represent documents in both a formatted
3131form and a processable (ie revisable) form, thus allowing
3132both the content and the printed appearance of the
3133document to be unambiguously transferred.
3134
3135In addition to text data, ODA supports graphics and image
3136data. A revised version to be published in 1993 will
3137support colour. Future developments include support for
3138audio content (underway) and video content (planned). An
3139interface to MHEG is also planned.
3140
3141ODA differs from SGML in that the former concerns itself
3142with the physical appearance of the document, while SGML
3143deliberately avoids doing so. SGML concerns itself with
3144semantic markup, and can be used to describe a wide range
3145of data and document architectures. ODA has a more
3146limited concept of a document.
3147
3148Hypermedia extensions to ODA (HyperODA) are underway.
3149The extensions will support:
3150
3151 References to data held externally to the document
3152 (similar to SGML's external entities?).
3153
3154 Non-linear structures, using contextual and independent
3155 hyperlinks based on the HyTime model.
3156
3157 Temporal relationships between document components (eg
3158 sequential, parallel, cyclic, duration, start delay).
3159
3160HyperODA is not being developed in competition to HyTime
3161or MHEG - its purpose is to add hypermedia features to
3162ODA rather than to be a completely general framework for
3163hypermedia applications.
3164
3165Bearing in mind that:
3166
3167 the HyperODA extensions are still under development;
3168
3169 in some senses ODA can be seen as a competitor to SGML,
3170 which has greater presence in the hypermedia world;
3171
3172 there seems to be a lack of enthusiasm for ODA in the
3173 Internet community (the IETF WG on piloting ODA has
3174 disbanded);
3175
3176 Adobe's newly-released Acrobat technology (described
3177 below) will have a significant effect on the
3178 marketplace;
3179
3180this report recommends that ODA should not form a basis
3181for investment in networked hypermedia technology by
3182RARE.
3183
3184 PREMO
3185
3186PREMO (Presentation Environment for Multimedia Objects)
3187is a new work item in ISO/IEC JTC1/SC24 (the graphics
3188standards subcommittee). An initial draft [ISO92a]
3189exists, and the schedule calls for a CD by June 1994, a
3190DIS by June 1995, and the final IS by June 1996.
3191
3192PREMO addresses the construction of, presentation of, and
3193interaction with multimedia objects. It specifies
3194techniques for creating audio-visual interactive single
3195and multiple media applications. It is consistent with
3196the principles of the Computer Graphics Reference Model
3197(CGRM, ISO 11072), and is defined in object-oriented
3198terms.
3199
3200It is not clear how PREMO relates to HyTime and MHEG.
3201Although these standards are listed in section 2
3202(References) of the initial draft, they appear not to be
3203mentioned in the text. The wisdom of developing what
3204appears to be yet another structuring standard for
3205multimedia data is doubtful.
3206
3207The PREMO work is not sufficiently advanced to permit a
3208judgement of its usefulness in satisfying the
3209requirements under discussion.
3210
3211 Acrobat
3212
3213Adobe Inc has introduced a new format called Acrobat PDF,
3214which it is putting forward as a potential de facto
3215standard for portable document representation. Based on
3216the Postscript page description language, Acrobat PDF is
3217also designed to represent the printed appearance of a
3218document (which may include graphics and images as well
3219as text. Unlike postscript however, Acrobat PDF allows
3220data to be extracted from the document. It is thus a
3221revisable format. It includes support for annotations,
3222hypertext links, bookmarks and structured documents in
3223markup languages such as SGML. PDF files can represent
3224both the logical and the formatting structure of the
3225document.
3226
3227Acrobat PFD thus appears to offer very similar
3228functionality to ODA. Adobe's successful Postscript de
3229facto standard profoundly influenced information
3230technology - it is possible that if successful, Acrobat
3231PDF will be almost as important. RARE should be aware of
3232this technology and its potential impact on multimedia
3233information systems.
3234
3235 Access Mechanisms
3236
3237This section describes some standards which are useful in
3238providing network access to multimedia data. Of course,
3239there are many multimedia transport protocols, which this
3240report does not attempt to describe (see [Adi93] for
3241further information). The protocols mentioned below are
3242search/retrieve protocols which were not mentioned in
3243[Adi93].
3244
3245 Multimedia Extensions to SQL
3246
3247A new work item in ISO (ISO/IEC JTC1 N2265) to extend the
3248SQL standard to include multimedia data is expected to be
3249approved shortly. Initially this work will concentrate
3250on developing a framework, and on free text data.
3251Support for non-text data will be added later, using a
3252separate part of the standard for each media type.
3253
3254The expected timescale for this standardisation work is
3255lengthy (part 1 - the framework - is targeted for
3256completion in 1996).
3257
3258There are suggestions that this standard could be used as
3259a query language in conjunction with the HyQ query
3260component of the HyTime standard.
3261
3262 DFR
3263
3264DFR is the Document Filing and Retrieval system,
3265specified in ISO 10166-1 and ISO 10166-2. It is intended
3266for office automation applications, and falls within the
3267Distributed Office Applications (DOA) model of ISO 10031-
32681. DFR has design similarities to the ISO Directory and
3269to the X.400 Message Store, and it is likewise part of
3270OSI.
3271
3272DFR defines a Document Store, which provides a service to
3273a DFR User over an OSI protocol stack incorporating ROSE
3274(and optionally RTSE). A document in the Document Store
3275may have a number of attributes associated with it,
3276including pointers to related documents. There is
3277support for multiple versions of the same document, and
3278for hierarchical groups of documents. The access
3279protocol supports searching for documents based on their
3280attributes. DFR itself does not restrict the content of
3281documents in any way, but the natural partner to DFR is
3282the ODA standard for document content.
3283
3284It is not clear that DFR offers significantly more useful
3285functionality than is available from other, simpler
3286access protocols already in use on the Internet.
3287
3288 Other Standards
3289
3290This section briefly describes other standards in this
3291area and discusses their relevance.
3292
3293 MIME
3294
3295MIME (Multipurpose Internet Mail Extensions) is a
3296mechanism for transferring multimedia information in an
3297RFC822 mail message. RFC 822 defines a message
3298representation protocol which specifies considerable
3299detail about message headers, but which leaves the
3300message content as flat ASCII text. RFC1341 redefines
3301the format of message bodies to allow multi-part textual
3302and non-textual message bodies to be represented and
3303exchanged without loss of information. Because RFC 822
3304said very little about message content, RFC 1341 is
3305largely orthogonal to (rather than a revision of) RFC
3306822.
3307
3308MIME provides facilities to include multiple objects in a
3309single message, to represent text in character sets other
3310than US-ASCII, to represent formatted multi-font text
3311messages, to represent non-textual material such as
3312images and audio fragments, and generally to facilitate
3313later extensions defining new types of Internet mail for
3314use by co-operating mail agents. It does not define any
3315structure to allow relationships between body parts
3316within a message to be expressed.
3317
3318For the purposes of the requirements considered by this
3319report, the relevance of MIME is that it separates media
3320type from media encoding, and that it defines a procedure
3321for registering values of these attributes.
3322
3323The MIME construct of chief interest is the "Content-
3324Type" field. This contains a MIME "type" and "subtype",
3325and any "parameters" which further qualify the subtype.
3326The register of MIME content-types is maintained by the
3327Internet Assigned Numbers Authority (IANA). Content
3328types defined in the MIME standard itself include:
3329
3330Type Subtype Parameters Meaning
3331
3332text plain charset Plain text
3333
3334 richtext charset Text with SGML-
3335 like markup for
3336 representing
3337 formatting.
3338
3339image jpeg JPEG File
3340 Interchange Format
3341
3342 gif Graphics
3343 Interchange Format
3344
3345audio basic 8-bit -law 8kHz
3346 PCM encoding
3347
3348video mpeg
3349
3350applicati ODA profile Open Document
3351on (used (Document Architecture
3352for Application document.
3353applicati Profile)
3354on-
3355specific
3356data)
3357
3358 octet- name (eg General binary
3359 stream filename); data such as an
3360 type (for arbitrary binary
3361 human file.
3362 recipient),
3363 etc.
3364
3365 postscri Document in
3366 pt postscript.
3367
3368
3369
3370Private experimental values of types and subtypes
3371starting with X- may be used between consenting adults
3372without registration with IANA.
3373
3374MIME also defines a "Content-Transfer-Encoding" field,
3375which is used to specify an invertible mapping between
3376the "native" encoding of a media type and a
3377representation that may be readily exchanged using 7-bit
3378mail transfer protocols.
3379
3380WWW's HTTP2 protocol makes use of MIME media type and
3381encoding attributes, and also uses MIME's message format
3382for retrieving data from the server. It is the first
3383MIME application to utilise the 8bit Content-Transfer-
3384Encoding, which essentially means no encoding.
3385
3386 SMSL
3387
3388SMSL is the Standard Multimedia Scripting Language. It
3389is a proposed new work item for ISO/IEC JTC1/SC18/WG8
3390(HyTime) and JTC1/SC29/WG12 (MHEG). The functional
3391requirements are expected to be completed in 1994, and
3392the coding scheme completed in 1995.
3393
3394SMSL is designed as an open language with a similar
3395purpose to existing vendor-specific scripting languages
3396such as Macromind's "Lingo", Kaleida's "Script/X", and
3397Gain's "GEL". The intention is to offer an intermediate
3398open multimedia scripting language which could be used
3399both for interchange purposes, and for controlling the
3400presentation of HyTime or MHEG multimedia structures.
3401Several different approaches to defining SMSL have been
3402suggested, including using the ANDF (Architecture-Neutral
3403Distribution Format) approach, and basing SMSL on SGML or
3404on the Scheme language.
3405
3406The SMSL work is not sufficiently advanced to permit a
3407judgement of its usefulness in satisfying the
3408requirements under discussion. However, it is
3409interesting to note that despite the descriptive power of
3410HyTime and MHEG, there is still perceived to be a role
3411for procedural scripting.
3412
3413 AVIs
3414
3415The CCITT is defining a set of Audio Visual Interactive
3416Services (AVIs), intended for offering to domestic and
3417business consumers over a national network (eg by PTTs).
3418These services will be specified as T.17x
3419recommendations, and will include MHEG. These services
3420would also make use of the SMSL work.
3421
3422Insufficient information is available about this area to
3423allow its relevance to be judged.
3424
3425 Trade Associations
3426
3427Thos section mentions some trade associations which are
3428involved in standards making in the multimedia area.
3429
3430 Interactive Multimedia Association
3431
3432The Interactive Multimedia Association (IMA) is an
3433international trade association with over 250 members,
3434representing a wide spectrum of multimedia industry
3435players. Members include Apple, Microsoft, MIT CECI (the
3436developers of AthenaMuse 2), 3DO, and many other
3437important market actors.
3438
3439In 1989, the IMA initiated a "Compatibility Project",
3440tasked with developing technical solutions to the cross-
3441platform compatibility problem. The Project has
3442published two important documents:
3443
3444 "Recommended Practices for Multimedia Portability"
3445 [IMA90] outlines a specification for a common interface
3446 to be used by interactive video delivery systems. It
3447 has been adopted by the US Military as part of Military
3448 Standard 1379.
3449
3450 "Recommended Practices for Enhancing Digital Audio
3451 Compatibility in Multimedia Systems" [IMA92] defines
3452 four standard digital audio data types and four
3453 sampling rates (from low-end -law 8kHz mono encoding,
3454 up through ADPCM modes to CD-quality 44kHz 16-bit
3455 stereo).
3456
3457Work is continuing to produce further recommendations on
3458other issues.
3459
3460The Compatibility Project has now initiated a procurement
3461process by publishing three Request for Technology (RFT)
3462documents, defining the requirements of a platform-
3463independent interactive multimedia system, including
3464networking requirements. The RFTs cover "Multimedia
3465System Services", a "Scripting Language for Interactive
3466Multimedia Titles", and "Multimedia Data Exchange". An
3467"Architecture Reference Model" for cross-platform desktop
3468and distributed multimedia systems provides the framework
3469for these RFTs, which are pragmatic documents outlining
3470the technical requirements for time-based media handling
3471in detail. Note that relatively little is said about non-
3472time-based data.
3473
3474A first reading of the Multimedia Data Exchange RFT
3475reveals that the Apple Bento standard [HR92] and the
3476Microsoft/IBM RIFF format [Mic92] both influenced the
3477development of this document. The selected system may
3478well be based on one or both of these technologies.
3479
3480A joint response to the Multimedia System Services RFT
3481has been received from HP, IBM and Sun. Two responses to
3482the Scripting Languages RFT have been received - from
3483Kaleida (Script-X) and Gain Technology (GEL). Two
3484partial responses to the Multimedia Data Exchange RFT
3485have been received from Apple (Bento) and Avid (Open
3486Media Framework).
3487
3488Responses to the RFTs are currently being analysed by the
3489IMA, and the result will be announced in November 1993.
3490The specifications which will eventually result from this
3491process will be important for future commercial
3492multimedia products. It is important that the community
3493keep a watching brief on the IMA Compatibility Project
3494and its possible implications for distributed multimedia
3495applications on the Internet.
3496
3497 Multimedia Communications Forum
3498
3499The Multi-Media [sic] Communications Forum (MMCF) is a
3500recently-formed (June 1993) trade consortium whose
3501initial members include IBM, National Semiconductor,
3502Apple, Siemens and AT&T. Intended to complement the work
3503of the IMA, the MMCF plans to develop guidelines and
3504recommendations for the industry to help ensure "end-to-
3505end network interconnectivity of multimedia applications,
3506workstations and devices". They also plan to provide
3507input to standards bodies.
3508
3509It is still too early to say whether this forum will
3510succeed. If the IMA Compatibility Project
3511specifications, when they are published, leave networking
3512issues open, then MMCF could have an important role to
3513play. It is recommended that RARE consider becoming an
3514Observing Member ($350 US pa), entitling it to attend
3515general and annual MMCF meetings (but not committee
3516meetings), and to receive minutes and other general
3517papers (but not working documents); with the prospect of
3518becoming an Auditing Member ($1200 US pa) later if
3519relevant.
3520
3521Multimedia Communications Community of Interest
3522
3523This is a very new organisation formed at a meeting in
3524France in June 1993. Its charter is to promote the use
3525of applications which let people in different locations
3526view documents, images, graphics and full-motion video on
3527a PC screen. The remit includes CSCW aspects. Members
3528of the organisation include IBM, Intel, Northern Telecom,
3529Telstra (Australia), BT, France Telecom and DB Telekom.
3530The companies plan field trials of multimedia services in
35311Q94.
3532
3533 Future Directions
3534
3535 General Comments on the State-of-the-Art
3536
3537Distributed hypermedia systems are now emerging from the
3538research phase into the experimental deployment stage.
3539Every project team (and standards committee), almost
3540without exception, hopes for their system to become the
3541de facto standard for hypermedia.
3542
3543As we've seen, Gopher and WWW already offer multimedia
3544capability, but they are still largely oriented to the
3545use of external viewers for non-text nodes. This
3546"unintegrated" approach is in contrast to typical stand-
3547alone multimedia applications, where the presentation of
3548related information in different media is tightly
3549integrated. The in-line image feature of XMosaic and the
3550new version of HTML currently under development may
3551represent the start of a move towards greater integration
3552of different media in such distributed hypermedia
3553systems.
3554
3555Three important factors in the design of distributed
3556hypermedia systems appear to emerge from the preceding
3557chapters of this report. They can each be formulated in
3558terms of distinctions between two aspects of the system.
3559
3560 A common and apparently fruitful approach to hypermedia
3561 systems is to distinguish the content from the
3562 hyperstructure. Standards work clearly distinguishes
3563 between these concepts, with standards such as MPEG,
3564 JPEG, G.72x, etc, for content; and HyTime or MHEG for
3565 structure. Currently-deployed systems also make this
3566 distinction, most obviously in Gopher, where the
3567 structure/content split maps onto the server
3568 filesystem's directory/file split. In a similar way,
3569 the ability to maintain hyperlink information
3570 separately from data is perceived in hypermedia
3571 research circles as a "good thing". Research systems
3572 such as Microcosm and Hyper-G do this, and HyTime with
3573 its ilink element also supports it. WWW does not
3574 support this, but requires link anchors to be edited
3575 into source data. There are problems with this
3576 approach, however - see the section on Microcosm for
3577 details.
3578
3579 A useful approach to content is to distinguish the
3580 media type from the media encoding. The MIME standard
3581 (used by HTTP2) illustrates how this can be done, and
3582 Gopher+ employs a similar system.
3583
3584 The distinction between data and protocol is also
3585 important for some systems. WWW for instance has
3586 clearly separate protocol (HTTP) and data (HTML)
3587 specifications. However, Gopher+ is specified without
3588 making this distinction. (The original Gopher system
3589 is very simple and arguably has no need for such
3590 separation.)
3591
3592The most significant mismatches between the capabilities
3593of currently-deployed systems and user requirements are
3594in the areas of presentation and quality of service.
3595Adding flexibility in presentation capabilities to WWW or
3596Gopher should be possible without any major change to the
3597protocols (although it may require changes to data
3598formats). Such capabilities could result from the
3599progress towards greater integration of media types
3600presaged above. However, improving QOS is significantly
3601more difficult, as it may require changes at a more
3602fundamental level. The following section outlines some
3603possible solutions to this problem.
3604
3605 Quality of Service
3606
3607Meeting the responsiveness requirement is certainly the
3608key factor for the acceptance of networked multimedia
3609information systems in the user community. To reiterate
3610the requirement given in a previous section:
3611
3612 For simple actions such as "next page", tolerable
3613 delays are of the order of 0.2s.
3614
3615 For more complex actions such as "search for documents
3616 containing this word", then a tolerable delay is of the
3617 order of 2s.
3618
3619 Users tend to give up waiting for a response after
3620 about 20s.
3621
3622There are several methods which may alleviate the problem
3623of poor responsiveness (or cause the user to revise his
3624or her expectations of responsiveness!), some of which
3625are described below.
3626
36271.Give clues that fetching a particular item might be
3628 time-consuming - simply quoting the size (and/or
3629 location) may be sufficient. WAIS and some Gopher
3630 clients already quote the size.
3631
36322.Display a "progress" indicator while fetching data.
3633
36343.Allow the user to interact with other, previously
3635 fetched information while waiting for data to be
3636 retrieved. The inability to do this is an annoying
3637 limitation of XMosaic. It can be difficult to
3638 implement, except on a multi-threaded operating system
3639 such as OS/2 or Windows NT.
3640
36414.Allow several fetches to be performed in parallel.
3642 Again, multithreading support makes this easier. This
3643 technique is less likely to be useful if all the nodes
3644 being requested come from the same server.
3645
36465.Pre-fetch information which the client software
3647 believes the user will wish to see next. This requires
3648 some "hints" in the data about which nodes might be
3649 good candidates for pre-fetching.
3650
36516.Cache information locally. The use of Universal
3652 Resource Numbers (see the section on WWW) is relevant
3653 for managing this.
3654
36557.Where multiple copies of the same information are held
3656 in different network locations, fetch the "nearest"
3657 copy. This is sometimes known as "anycasting", and is
3658 a more general case of local caching. The proposed URN-
3659 to-URL resolution service [WD93] could be used to
3660 support this.
3661
36628.When retrieving a document, the client should be able
3663 to display the first part of the document to the user.
3664 The user can then start to read the document while the
3665 system is still downloading it. Alternatively, the
3666 user may decide that the document is not relevant and
3667 abort the retrieval.
3668
36699.Offer multiple views of image or video data at
3670 different resolutions and therefore sizes. This
3671 enables the user to select a balance between speed of
3672 retrieval and data quality. Gopher+ and HTML2 both
3673 support this.
3674
367510. Future high-speed networks and protocols (ATM,
3676 RTP) will allow real-time display of isochronous data.
3677 Information systems should be able to take advantage of
3678 this.
3679
3680A useful description of the problem is given in [Loe92].
3681This paper rightly contends that the view, held by many
3682hypermedia researchers and implementors, that the network
3683is simply a transparent data highway which needs no
3684special consideration in application design, is wrong.
3685It is argued that:
3686
3687 "the very same structural characteristics that may
3688 make a multimedia document appealing to the end user
3689 are the characteristics that are extremely helpful
3690 during dynamic network performance optimisation".
3691
3692This is a particularly relevant statement considered in
3693the light of suggestion 5 above.
3694
3695 Recommended Further Work
3696
3697To meet the needs of applications such as those described
3698in section 2.1, the community must seek where possible to
3699adapt and enhance existing tools, not to build new ones.
3700There is now an opportunity for RARE to stimulate and
3701encourage this process of adaptation and enhancement, and
3702the following subsections outline a strategy for this.
3703
3704 Selecting a System
3705
3706In order to have the greatest effect, RARE should
3707concentrate its efforts on only one of the existing
3708tools. Candidate technologies are those already
3709outlined: Gopher, WWW, WAIS, Hyper-G, Microcosm and
3710AthenaMuse 2.
3711
3712It is recommended that RARE should select the World-Wide
3713Web to concentrate its efforts on. The reasons for this
3714decision are as follows.
3715
3716 Flexibility. The rich yet straightforward design of
3717 WWW, with its clearly separable components (HTML, URL
3718 and HTTP), means that it is a very flexible basis on
3719 which to develop distributed multimedia applications.
3720
3721 Existing efforts. The WWW implementor community is
3722 already discussing and designing extensions to HTML
3723 (HTML2), intended (among other things) to support
3724 multimedia. There is clearly much interest in this
3725 area, and RARE efforts could complement existing work.
3726
3727 Hyperlinks. A clear requirement of many applications
3728 is the availability of hyperlinking, which WWW supports
3729 well.
3730
3731 Integrated solution. Because WAIS, Gopher and Hyper-G
3732 (as well as anonymous FTP servers) may all be accessed
3733 from Web clients, WWW serves as an important
3734 integrating tool for information services. It is
3735 important that distributed multimedia applications,
3736 which require extensive support in the client software,
3737 should be based on a technology "close to" such
3738 integrated clients.
3739
3740 Penetration and growth. Although Gopher far surpasses
3741 WWW in the number of servers available, the rate of
3742 growth in WWW usage is greater than that of Gopher.
3743 There is an increasing realisation in the community
3744 that Gopher is over-simplistic for many purposes, and a
3745 corresponding increase in interest in WWW.
3746
3747 Attention to QOS issues. There is already an awareness
3748 in the WWW community of the need for achieving an
3749 appropriate QOS, and a mechanism has already been
3750 proposed in HTTP2 to alleviate the problem.
3751
3752 Standardisation. The WWW team is taking
3753 standardisation of the existing WWW system components
3754 seriously. The URL format has already been published
3755 as an Internet draft (and has been adopted as an
3756 important component of the proposed Internet integrated
3757 information infrastructure), and the current version of
3758 HTML is about to follow suit. The use of SGML as the
3759 basis of HTML complies with the perceived importance of
3760 SGML for hypermedia in general (and also fits in with
3761 RARE's approach of adopting appropriate open
3762 standards).
3763
3764 Software status. CERN has recently placed the WWW code
3765 developed by it into the public domain. This is unlike
3766 all the other candidate technologies, which all have
3767 restrictions on who can do what with the code. In the
3768 case of Gopher, these restrictions are already causing
3769 some commercial users to look at other options.
3770
3771WWW has two significant disadvantages, both of which are
3772being alleviated:
3773
3774 Restricted choice of client software. At present,
3775 Apple Macintosh and PC/MS Windows clients are available
3776 in beta form only. By contrast, there are more than
3777 one well-tested Gopher clients available for these
3778 platforms.
3779
3780 However, other WWW clients for the Mac and MS Windows
3781 are in the pipeline.
3782
3783 There is a perception in the community that making
3784 information available over HTTP is difficult, and that
3785 it must be put into HTML.
3786
3787 However, it is possible to put plain-text, non-HTML
3788 documents onto the Web. Such documents of course
3789 cannot contain links. Furthermore, WYSIWYG HTML text
3790 editors are available, to ease the pain of writing
3791 HTML.
3792
3793The main disadvantages of the other systems are:
3794
3795 Gopher is designed for simplicity, and therefore lacks
3796 the flexibility of WWW. In particular its structure is
3797 too inflexibly hierarchical and it does not have
3798 hyperlinks. Its main advantage is its very heavy
3799 penetration. However, because of the WWW approach to
3800 accessing data using other protocols, all of
3801 gopherspace is part of the Web. Any Web client should
3802 be able to be a gopher client too.
3803
3804 It is neither envisaged that Gopher will go away, nor
3805 that it won't be used for multimedia data. However,
3806 Gopher is unlikely to be used for more sophisticated
3807 multimedia applications such as academic publishing,
3808 interactive multimedia databases and CAL, because of
3809 the above-mentioned limitations.
3810
3811 WAIS is a specialised tool, and will certainly form
3812 part of the overall solution, particularly for database-
3813 type applications. It is not a general solution for
3814 distributed hypermedia applications.
3815
3816 AthenaMuse 2 is commercially-oriented: it is clear that
3817 academic and research users will have to pay to use the
3818 software. Its level of use is thus very unlikely to be
3819 as great as publicly-available systems such as WWW.
3820 Moreover, it does not support all the required
3821 platforms.
3822
3823 Microcosm network support is still in early stages,
3824 limited at present to the PC/Windows platform. If it
3825 can be shown to perform adequately over a network, if
3826 it is capable of scaling to global levels, and if the
3827 advantages of maintaining link information separately
3828 from documents are found clearly to outweigh the
3829 consequent difficulties, it may become important in the
3830 future. Microcosm's authors need to ensure that the
3831 commercialisation of Microcosm does not hinder its
3832 adoption by the academic community.
3833
3834 Hyper-G is more difficult to dismiss. It is still in a
3835 relatively early stage of development, but appears to
3836 have many of the necessary features. Its main
3837 disadvantages are: (a) the lack of penetration outside
3838 the University of Graz - the author is aware of only
3839 one other site using it; and (b) it is currently
3840 limited to UNIX only. The author believes that, given
3841 WWW's head start in terms of deployment, and the
3842 current progress in adding multimedia facilities to it,
3843 WWW stands a much better chance than Hyper-G of being
3844 accepted as the de facto standard for distributed
3845 multimedia applications on the Internet.
3846
3847 Directions for RARE
3848
3849Earlier in this report, it was noted that the most
3850important areas where effort was needed were (a)
3851provision of facilities for the integrated presentation
3852of multimedia data (including synchronisation issues);
3853and (b) ensuring adequate responsiveness.
3854
3855Bearing this in mind, it is recommended that RARE should
3856invite proposals and (subject to funding being available)
3857subsequently commission work to:
3858
38591.Develop conversion tools from commercial authoring
3860 packages to WWW, and establish authoring guidelines for
3861 authors who wish to use the conversion tools. This is
3862 a significant and high-profile development aimed at
3863 enabling sophisticated multimedia applications to run
3864 over the network. (Authoring guidelines will be
3865 necessary to enable authors to fit in with the Web's
3866 way of doing things, and to document features of the
3867 authoring package which should be avoided because of
3868 conversion difficulties.)
3869
38702.Implement and evaluate the most promising ways of
3871 overcoming the QOS problem. This is an essential task
3872 without which interactive distributed multimedia
3873 applications cannot become a reality. Some
3874 possibilities have already been outlined in the
3875 preceding chapter.
3876
38773.Implement a specific user project using these tools, in
3878 order to validate that the facilities being developed
3879 are truly relevant to actual user requirements. It may
3880 be that partner funding from the selected user project
3881 would be appropriate.
3882
38834.Use the experience gained from 1, 2 and 3 to inform and
3884 influence the further development of HTML2 and HTTP2 to
3885 ensure that they provide the required facilities.
3886
38875.Contribute to the development of the WWW clients
3888 (particularly the Apple Macintosh and PC/MS Windows
3889 clients) in terms of their multimedia data handling
3890 facilities.
3891
3892Although it is strictly speaking outside the remit of
3893this report (since it is not specifically concerned with
3894multimedia data), it is noted that the rapid growth of
3895WWW may in the future lead to problems through the
3896implementation of multiple, uncoordinated and mutually
3897incompatible add-on features. To guard against this
3898trend, it may be appropriate for RARE, in coordination
3899with CERN and other interested parties such as NCSA, to:
3900
39016.Encourage the formation of a consortium to coordinate
3902 WWW technical development (protocol enhancements, etc).
3903
3904 References
3905
3906[Adi93] "A Survey of Distributed Multimedia
3907 Research, Standards and Products", ed. C
3908 Adie, January 1993 (RARE Technical Report
3909 5). URL=ftp://ftp.ed.ac.uk/pub/mmsurvey/
3910
3911[AAL92] "Gopher+: Proposed Enhancements to the
3912 Internet Gopher Protocol", B Alberti, F
3913 Anklesaria, P Linder, M McCahill, D
3914 Torrey, Summer 1992.
3915 URL=gopher://boombox.micro.umn.edu:70/11/g
3916 opher/gopher_protocol/Gopher%2b
3917
3918[BCH92] "The AthenaMuse 2 Functional
3919 Specification", L Bolduc, J Culbert T
3920 Harada, J Harward, E Schlusselberg, May
3921 1992.
3922 URL=ftp://ceci.mit.edu/pub/AM2/funcspec.tx
3923 t.Z
3924
3925[BF92] "MIME (Multipurpose Internet Mail
3926 Extensions): Mechanisms for Specifying and
3927 Describing the Format of Internet Message
3928 Bodies", RFC1341, N. Borenstein & N.
3929 Freed.
3930 URL=ftp://src.doc.ic.ac.uk/rfc/rfc1341.txt
3931 .Z
3932
3933[CEC92a] "Research and Technology Development in
3934 Advanced Communications Technologies in
3935 Europe: RACE '92", CEC, March 1992.
3936 Available from: raco@postman.dg13.cec.be
3937
3938[CEC92b] "Esprit Programme Synopses", CEC, October
3939 1992. In seven volumes. Available from
3940 esprit_order_mailbox@eurokom.ie
3941
3942[DHH92] "Towards an Integrated Information
3943 Environment with Open Hypermedia Systems",
3944 H Davis, W Hall, I Heath, G Hill,
3945 Proceedings of the ACM Conference on
3946 Hypertext, Milan 1992, p181-190.
3947
3948[DKM90] "WAIS Interface Protocol", F Davies, B
3949 Kahle, H Morris, J Salem, T Shen, R Wang,
3950 J Sui and M Grinbaum, April 1990.
3951 URL=ftp://quake.think.com/wais/doc/protspe
3952 c.txt
3953
3954[Eli93] "Deja-Vu Distributed Hypermedia
3955 Application Framework", A Eliens.
3956 URL=ftp://ftp.cs.vu.nl/eliens/Deja-Vu-
3957 proposal.ps
3958
3959[FBD93] "A Status Report on Networked Information
3960 Retrieval: Tools and Groups", ed. J
3961 Foster, G Brett and P Deutsch, March 1993.
3962 URL=ftp://mailbase.ac.uk/pub/nir/nir.statu
3963 s.report
3964
3965[ISO92a] "Initial Draft PREMO (Presentation
3966 Environment for Multimedia Objects",
3967 ISO/IEC JTC1/SC24 N847, November 1992.
3968
3969[HBR93] "The Amsterdam Hypermedia Model: extending
3970 hypertext to support real multimedia", L
3971 Hardman, D C A Bulterman, G van Rossum,
3972 Amsterdam 1993
3973 URL=ftp://ftp.cwi.nl/pub/CWIreports/CST/CS-
3974 R9306.ps.Z
3975
3976[HR92] "Bento Specification", J Harris and I
3977 Ruben, Apple Computer Inc, August 1992.
3978 URL=ftp://ftp.apple.com/apple/standards/Be
3979 nto_1.0d4.1
3980
3981[HS90] "The Dexter Hypertext Reference Model", F
3982 Halasz and M Schwartz, NIST Hypertext
3983 Standardisation Workshop, January 1990.
3984
3985[IMA90] "Recommended Practices for Multimedia
3986 Portability", Release 1.1 October 1990,
3987 Interactive Multimedia Association, 3
3988 Church Circle, Suite 800, Annapolis, MD
3989 21401-1993, USA.
3990
3991[IMA92] "Recommended Practices for Enhancing
3992 Digital Audio Compatability in Multimedia
3993 Systems", Release 3.00 1992, Interactive
3994 Multimedia Association, 3 Church Circle,
3995 Suite 800, Annapolis, MD 21401-1993, USA.
3996
3997[Kim93] Article in comp.text.sgml newsgroup, 24
3998 May 1993, by Eliot Kimber
3999 (drmacro@vnet.ibm.com).
4000 URL=ftp://ftp.ifi.uio.no/SGML/comp.text.sg
4001 ml/by.msgid/19930524.152345.29@almaden.ibm
4002 .com
4003
4004[KS92] "Hyper-G: A Universal Hypermedia System",
4005 F Kappe and N Sherbakov, March 1992.
4006 URL=ftp://iicm.tu-graz.ac.at/pub/Hyper-
4007 G/doc/report333.txt.Z
4008
4009[Lee92] "The HTTP Protocol as Implemented in W3",
4010 T Berners-Lee, January 1992.
4011 URL=ftp://info.cern.ch/pub/www/doc/http.tx
4012 t
4013
4014[Lee93a] "Uniform Resource Locators", T Berners-
4015 Lee, March 1993.
4016 URL=ftp://info.cern.ch/pub/ietf/url4.ps
4017
4018[Lee93b] "Protocol for the Retrieval and
4019 Manipulation of Textual and Hypermedia
4020 Information", T Berners-Lee, 1993.
4021 URL=ftp://info.cern.ch/pub/www/doc/http-
4022 spec.ps
4023
4024[Lee93c] "Hypertext Markup Language (HTML)", T
4025 Berners-Lee, March 1993.
4026 URL=ftp://info.cern.ch/pub/www/doc/html-
4027 spec.ps
4028
4029[Loe92] "Delivering Interactive Multimedia
4030 Documents over Networks", S Loeb, IEEE
4031 Communications Magazine, May 1992.
4032
4033[Mar91] "Emerging Hypermedia Standards" B Markey,
4034 Multimedia for Now and the Future (Usenix
4035 Conference Proceedings), June 1991.
4036
4037[Mic92] "RIFF Tagged File Format", Microsoft Inc,
4038 1992.
4039
4040[NN93] "Davenport Advisory Standard for
4041 Hypermedia (DASH), Module I: Standard Open
4042 Formal Architecture for Browsable
4043 Hypermedia Documents (SOFABED)", ed S R
4044 Newcomb and V T Newcomb.
4045 URL=ftp://sgml1.ex.ac.uk/davenport/sofabed
4046 .0.9.6.ps.Z
4047
4048[RJM93] "CMIFed: A Presentation Environment for
4049 Portable Hypermedia Documents", G van
4050 Rossum, J Jansen, K S Mullender, D C A
4051 Bulterman, Amsterdam 1993 (also presented
4052 at ACM Multimedia 93 conference).
4053 URL=ftp://ftp.cwi.nl/pub/CWIreports/CST/CS-
4054 R9305.ps.Z
4055
4056[Shn84] "Response Time and Display Rate in Human
4057 Performance with Computers", B
4058 Shneiderman, Comp. Surveys 16, 1984.
4059
4060[Str91] "Directory of Electronic Journals and
4061 Newsletters", Edition 1 (July 1991), M
4062 Strangelove.
4063 URL=ftp://137.122.6.16/pub/religion/electr
4064 onic-serials-directory.txt
4065
4066[WD93] "A Vision of an Integrated Internet
4067 Information Service", C Weider and P
4068 Deutsch, March 1993.
4069 URL=ftp://ietf.cnri.reston.va.us/internet-
4070 drafts/draft-ietf-iiir-vision-00.txt