World's First Ultra High Definition Shoulder-Mount Camera

This is the world's first compact shoulder-mount Ultra High Definition camera. Developed by NHK, it uses a single-chip color imaging sensor to produce 33MP video.

By reducing the size and weight of the camera, the portability had been improved, making it more maneuverable than previous prototypes, so it can be used in a wide variety of shooting situations. This compact head can also be used with commercially available still camera lenses.

As the single-chip sensor uses a Bayer color filter array, where only one color component is acquired per pixel, researchers at NHK have also developed a high quality up-converter, which estimates the other two color components to convert the output into full resolution video.

Next, NHK will develop a camera control unit to perform signal processing specifically for this head. This will improve the picture quality and functionality of the camera.


By Don Kennedy, DigInfo TV

HbbTV on Brink of Global Expansion

HbbTV is now odds on to emerge from the cluster of interactive and hybrid TV standards to become the dominant platform uniting broadcast and connected TV.

Currently sweeping through continental Europe, and under trial in a number of other major countries including China, Japan and the U.S., it looks like the winning hybrid TV platform is emerging, partly through not being too ambitious and sticking clearly within defined boundaries. It leaves plenty of scope for apps vendors, broadcasters and pay TV operators to innovate around the platform and stamp their own distinctive flavor on their products or standards.

The closest to an actual HbbTV deployment outside Europe is in South Korea, where national broadcaster KBS is launching services based on the country’s OHTV (Open Hybrid Television), which is a separate development but almost identical technically and now likely to be aligned completely with HbbTV given its momentum in other countries. In Korea, the service is hybrid digital terrestrial and broadband, as is the case with many early HbbTV deployments in Europe.

HbbTV evolved in 2009 as a joint project between France and Germany, and it is those two countries along with Spain and the Netherlands that have made the early running with HbbTV deployments. In Germany, at least eight broadcasters, including public service broadcaster RTL and pan European media distributor ZDF, are now offering HbbTV apps over terrestrial or satellite networks.

Such apps can combine multiple delivery channels into one coherent service, or provide additional viewing options within a single channel. An example of the first of these use cases is German home shopping channel QVC, which is exploiting HbbTV to unite its various existing distribution channels including TV, Internet and mobile networks, within one coherent service allowing customers to search for products. An example of the second is German private broadcaster Vox, part of the RTL Group, which is using HbbTV for a cooking channel, allowing viewers to access different recipes for demonstration within a show by pressing the “red button.”

In France, public broadcaster France Télévisions has been leading the way and is using HbbTV to expand coverage of the French Open tennis tournament over the next two weeks. It will allow viewers to choose from a number of matches at a given time, again using red button functionality.

Spain, meanwhile, has agreed to adopt HbbTV as its system for connected TV, with pilots completed by broadcaster Mediaset España and Telco Telefonica. These involved Mediaset’s Telecinco importing content from Telefónica’s services, including Movistar Imagenio, Movistar Videoclub and Terra TV.

Similarly, Dutch broadcasters, including SBS, NPO and RTL, have agreed to use HbbTV as their standard for hybrid connectivity, while launches have just occurred or are imminent in Switzerland, Austria, the Czech Republic and Poland.

HbbTV has also won over the Nordic region, which originally was planning to build hybrid broadcast around the alternative Multimedia Home Platform (MHP) developed by the DVB, with the main original difference being that HbbTV is based on HTML while MHP is written in Java language. Germany originally went with MHP, but it failed miserably, partly because there was then little demand for interactive TV in the country.

Now, the Nordic region comprising Denmark, Finland, Iceland, Norway and Sweden, oddly also including Ireland, has replaced DVB-MHP with HbbTV as the common API for hybrid digital receivers within its NorDig digital TV specification. The stated reason is that HbbTV now has much wide market acceptance, with a range of TV applications and new hybrid services, and crucially HbbTV compatible receivers from a number of manufacturers such as Humax of South Korea. Also significant is that HbbTV software stacks are now incorporated in leading hybrid chip sets from Broadcom and Sigma Designs. Most major players in the connected TV arena are now members of the HbbTV Forum.

The success of HbbTV can be put down to three factors: its flexibility; foundation on existing standards that are being implemented anyway as part of OTT and IPTV deployments; and support from industry groups, notably the European Broadcast Union (EBU), whose influence extends outside the continent.

The EBU is hoping that the Olympics will give HbbTV a big lift, and has laid the ground by offering “white-label” HbbTV applications free of charge to its members. These apps are currently being customized and rebranded for deployment just ahead of the games, when they will deliver interactive services to peak audiences during the Olympics.

Equally crucial is the approval of the Open IPTV Forum (OIPF), which has emerged as the major body forging standards for OTT and general online video service delivery over unmanaged infrastructures, as well as IPTV within closed ‘walled garden’ networks.

The key component for OTT that has been adopted by HbbTV is the OIPF’s Open Internet Profile, based on its Declarative Application Environment (DAE), which is a browser for TVs with support for various presentation mechanisms including HTML4, HTML5, SVG, and CE-HTML.

“Its key component is the set of JavaScript objects which permit the manipulation of media for both content on demand as well as live streaming, interactions with local and remote storage, and control of adaptive streaming,” said Nilo Mitra, OIPF president. “This specification is reused by the HbbTV Consortium for providing catch-up services via broadcaster portals, and is now implemented in retail TVs sold by major manufacturers in several EU regions,” Mitra added, arguing that this was the only fully open specification for content delivery over unmanaged networks available today.

The OIPF specification includes a mechanism for adaptive delivery of MPEG-2 transport streams over HTTP, and this has been incorporated into one of the profiles of the newly just published ISO DASH specifications, according to Mitra. DASH is likely to become the standard mechanism for adaptive streaming, taking over from existing proprietary systems such as Microsoft Smooth Streaming, and possibly Apple’s HLS (HTTP Live Streaming).

Support for DASH streaming was an addition to HbbTV. But, at the outset, it had the right foundation by being built on two relevant and mature technology sets, one being web standards already included in web browsers for embedded devices, and the second being the Digital Storage Media Command and Control (DSM CC) specification already part of the MHEG-5 interactive platform adopted in several countries, including the UK and Australia.

Use of existing web standards provided the basis for broadband access, while DSM CC specified a common approach for interactive services uniting two way online delivery with one way broadcast. DSM CC, therefore, delivered the hybrid component, and although it is a complex set of technologies the basic principle is simple. The aim is to facilitate control over transport of audio and video streams within interactive services in both a bi-directional environment such as a cable TV or VOD system, and also a uni-directional service such as satellite or digital terrestrial.

The challenge was how to simulate interactivity in a one-way environment when there was no return path. The simple answer was to adopt a carousel approach — hence the name. As the receiver in a traditional broadcast environment has no return path and so cannot request specific files from a server, DSM CC periodically transmits every file, and the receiver then grabs the ones it wants as they pass by on the carousel. Of course, this is not particularly efficient since if a receiver misses a file it has to wait for it come round again. But, techniques have been developed and embodied in HbbTV to improve the interactive performance for one way broadcast services.

It will become clear during the rest of 2012 how well HbbTV does perform in a variety of hybrid environments including those involving one-way broadcast, with the Olympics providing useful feedback. But, HbbTV has still not convinced all its doubters, even in Europe. Italy is the notable outsider, having gone its own way by deploying MHP for hybrid services.

The UK is the other major odd one out, since its much delayed connected TV platform, YouView, took a different approach. However, the UK is itself a bit of a hybrid, since the Digital Technology Group (DTG) responsible for digital TV and particularly terrestrial standards in the UK developed an extension to MHEG-5 called the MHEG-5 Interaction Channel (MHEG-IC), allowing broadcast interactive services to be delivered via an IP connection. This made it more like HbbTV, and, subsequently, the DTG has fully endorsed HbbTV for Freeview DTT.

By Philip Hunter, Broadcast Engineering

NHK 33 Megapixel 120fps Ultra High Definition Imaging System

NHK, in conjunction with Shizuoka University, has developed an Ultra High Definition imaging system that outputs 33MP video at 120fps.

As Ultra High Definition broadcasts at full resolution are designed for large, wall sized displays, there is a possibility that fast moving subjects may not be clear when shot at 60fps, so the option of 120fps has been standardized for these situations.

To handle the sensor output of approximately 4 billion pixels per second with a data rate as high as 51.2Gbps, a faster analog-to-digital converter has been developed to process the data from the pixels, and then a high-speed output circuit distributes the resulting digital signals into 96 parallel channels.

This 1.5-inch CMOS sensor is smaller and uses less power when compared to conventional Ultra High Definition sensors, and it is also the world's first to support the full specifications of the Ultra High Definition standard.

From now on NHK plan to increase the light sensitivity of this Ultra High Definition sensor.


by Don Kennedy, DigInfo

NHK Hybridcast Making broadcast TV Interactive

Hybridcast is an infrastructure system being developed by NHK, with a view to commercial use in 2013. This system combines broadcasting with the Internet, to enable a variety of TV-centered services.

At this year's NHK Science & Technology Research Laboratories Open Day, NHK exhibited prototype receivers, developed together with manufacturers, as well as the service concept.


Source: DigInfo

DPP Unveils Digital Workflow Guide

The Digital Production Partnership (DPP) has unveiled a major industry report, The Bloodless Revolution: A Guide to Smoother Digital Workflows in Television. The report is the first published guidance on digital workflows to be issued on behalf of ITV, Channel 4 and the BBC. It seeks to help producers and suppliers achieve a smoother transition to fully digital production.

The new guide follows the publication of the DPP’s report on breaking down the barriers to digital production The Reluctant Revolution, September 2011. One of the claims made in the first report was that the pace of change in the industry was held back by a lack of commonly agreed ways of working. It observed that greater guidance is needed if the industry is to complete its move from tape-based to file-based production.

The DPP’s new report now provides such guidance. It sets out to identify the smoothest, most efficient digital workflows for use with currently available technology, while providing sufficient background information to help maintain a view of the wider production landscape. It also identifies opportunities for collaboration, cost saving, and better creative outcomes.

The guide sets out a clear high level workflow as a framework for providing information, guidance and direction to digital production workflows. The overall process has been broken into four steps: planning – which covers the process up to the point of shooting, including the different conventions and practices that need to be adopted right at the outset; rushes management – which looks at the capture and handling of content on location or in studio up to the point of rushes archive and management; post production – which goes from the ingest of material for editing through to completion of the master: and delivery – the production of masters for delivery to broadcasters, clients or the audience.

The report was commissioned by the DPP from industry analysts MediaSmiths International. Its starting point was the views and experiences offered by dozens of attendees from all over the UK at the DPP’s regular industry forums.

From the outset ‘The Bloodless Revolution’ acknowledges that programme makers have no desire to see their world reduced to a series of workflows. Many may feel that by over-describing the process, the magic of television production will be driven out.

But the report goes on to offer a user-friendly map by which to navigate the potentially complex processes of file-based production – and in so doing offers a guide that, while first appearing analytical, is actually liberating.

Source: Digital Production Partnership

ITU Sets Standards for Ultra HD

The London Olympics will provide the first live trials of Ultra High Definition TV (UHDTV) based on standards finally agreed last week by the International Telecommunications Union (ITU) after a decade of research and development in which the EBU was heavily involved.

The ITU has defined the UHDTV standards, 4K and 8K, as multiples of the existing 1080p1920 format defined in the ITU-R Rec. 709 standard. HD 1080p, at present often referred to as full HD, displays at a resolution of 1920 pixels wide by 1080 high in progressive scan, corresponding to a widescreen aspect ratio of 16:9. Various frame rates are supported including 24, 50 and 60.

4K is defined simply by doubling 1080p1920 in each direction to yield pictures with four times the spatial resolution, at 3840 pixels wide by 2160 high, which is 8 mega pixels. 8K then doubles up again to resolution 7680 wide by 4320 high, spatially 16 times 1080p, or 32 mega pixels. But, the ITU has also added support for a higher frame rate option of 120, which experiments have shown may be necessary for accurate portrayal of motion at these very high resolutions on large wall sized displays.

Without the corresponding increase in frame rate, there is a danger that UHDTV will display brilliant images of slow moving action, but then exhibit slight jerkiness for high speed shots in some sporting events for example.

The bandwidth implications of these standards will alarm some operators and broadcasters, for an 8K programme running at the full 120 frames per second would require 320 times the bit rate of current HD transmissions given that these are often not yet even 1080p, but usually either 720p or interlaced 1080i. But, as the EBU pointed out, these new standards are unlikely to start working their way into mainstream transmissions for the best part of a decade.

That will coincide with the introduction of new frameless displays in which picture size can vary, and that blend into the background when not in use. Some vendors of pay-TV software are already developing platforms in anticipation of Ultra HD delivery to such large screens. For example, UK-based conditional access and middleware vendor NDS has a platform called Surfaces that was first demonstrated delivering 4K UHDTV to large displays at IBC 2011 in Amsterdam.

Most broadcasters will first upgrade to full HD at 1080p, which will for now meet all quality expectations, certainly for screens up to 60 inches diameter. For example, in the UK, the Freeview HD platform used by the BBC and ITV has been specified to provide full HD capability.

But although UHDTV may be some years away for TV, the 4K version has already been adopted for digital cinematography and computer graphics, using slightly different resolutions than the new ITU standard. 4K is also supported by YouTube, the only video hosting service to do so, in a different version again, allowing uploading of 4K videos at a resolution of 4096 x 3072 pixels, or 12.6 megapixels.

By Philip Hunter, Broadcast Engineering

Microsoft Announces Support for MPEG-DASH in Microsoft Media Platform

Microsoft Media Platform will support MPEG-DASH, a recently ratified ISO/IEC standard for dynamic adaptive streaming over HTTP. Microsoft plans to support DASH and other open standards as part of an industry-wide initiative to establish reliable video delivery to Internet connected devices and enable true interoperability between adaptive streaming technologies from different vendors.

Much like Smooth Streaming, DASH uses Extensible Markup Language (XML) to describe media presentations in a manifest file which references media streams stored in ISO Base Media File Format. Combined with the standard HTTP protocol and existing Web content delivery networks, the DASH standard enables a better video experience for end users by automatically adapting to varying client and network conditions during playback.

Taking advantage of similarities between Smooth Streaming and DASH, Windows Azure Media Services will add support for DASH Live Profile later this year so that both Smooth Streaming and DASH devices can access the same live and on-demand video presentations using either manifest format. This will enable a smooth transition to DASH for millions of devices and services currently using Smooth Streaming.

In addition to server-side support, Microsoft will also add support for DASH to all its Smooth Streaming client development kits. The first step will be to enable DASH support in the Smooth Streaming Client for Silverlight, followed by support in Smooth Streaming Client SDKs for Windows 8, iOS, Xbox, Windows Phone and Smooth Streaming Client Porting Kit for embedded devices.

Microsoft is also contributing to W3C efforts to standardize adaptive streaming APIs in HTML5 so that DASH Web applications may also be written in HTML5 and ECMAScript (JavaScript) in the future without requiring browser plug-ins such as Silverlight and Flash to enable advanced streaming media scenarios.

Microsoft has contributed to the development of the DECE UltraViolet video format which enables download and adaptive streaming of premium movie and TV content; and to various international broadcast standards and consortia so that a common protected video format based on DECE Common File Format, MPEG Common Encryption, and MPEG-DASH specifications will be supported by all adaptive streaming services and devices to enable reliable interoperability for consumers, just like broadcast TV and DVD.

What Microsoft services will have DASH support?
Windows Azure Media Services will provide encoding, encryption, and streaming support for Application Profiles based on DASH “ISO Base Media Live Profile” this year. Both DASH manifests and Smooth Streaming manifests will be generated to allow the same media to be streamed by DASH clients and Smooth Streaming clients. The primary media format will conform to the PIFF 1.3 specification in addition to Live Profile, will include several features and constraints compatible with the DECE Common File Format, and may optionally include MPEG Common Encryption with PlayReady DRM support. Windows Azure Media Services will also be capable of live transformation to multiple streaming formats, including MPEG-2 Transport Streams for use with DASH M2TS Simple Profile manifests or M3U8 playlists.

What Microsoft client technologies will have DASH support?
Microsoft plans to add MPEG-DASH support to all client development kits that currently support Smooth Streaming. These are: Smooth Streaming Client for Silverlight; Smooth Streaming Client for Windows Phone; Smooth Streaming Client SDK for Windows 8 Metro-style applications; Xbox LIVE Application Development Kit; Smooth Streaming SDK for iOS Devices with PlayReady; and Smooth Streaming Client Porting Kit.

Is Microsoft discontinuing Smooth Streaming?
No. Microsoft will continue to invest in Smooth Streaming as an established technology and brand while ensuring its Smooth Streaming services, clients, tools and workflows are DASH compatible. The Smooth Streaming file format (PIFF 1.3) is already compatible with the DASH specification (ISO Base Media Live Profile) so customers and partners who are investing into creation of Smooth Streaming content today will have a clear path to making that content deliverable to DASH clients in the future.

What is Common Encryption?
Common Encryption is an MPEG standard using AES-128 media encryption that enables a single protected ISO Base Media file or adaptive streaming presentation to be used with any DRM system supported by a device and the publisher. The standard is designated ISO/IEC 23001-7 “Information technology – MPEG systems technologies – Part 7: Common encryption in ISO base media file format files”. Prior to this standard, a different set of files was required for each different DRM type, and interchange of files between authorized devices was generally not possible because of different DRMs.

What is Common File Format?
Common File Format (CFF) is a DECE video specification titled “Common File Format & Media Formats Specification” used for content download. It specifies video files based on fragmented ISO Base Media files (MPEG-4 Part 12), optionally using Common Encryption, containing AVC video, AAC audio, SMPTE Timed Text and Graphics subtitles, metadata, and several optional audio formats. All parameters required for interoperability are sufficiently specified to allow independently implemented encoders, publishers, delivery services, and devices to reliably interchange and play the same files. Different “media profiles” are specified for high definition, standard definition, and “portable” definition devices.

The CFF requirement to use short movie fragments makes these files and compatible decoders forward compatible with DASH adaptive streaming using movie fragments as DASH Media Segments. DECE is currently in the process of specifying “Common Streaming Format” and considering DASH Application Profiles.

What about HTML5 playback?
The current working draft of HTML5 does not include specific support for either adaptive streaming or DRM protection. It is possible to indicate a playlist or manifest file as the source of the <video> tag, but a publisher would have no control over the behavior and presentation that each device or browser would execute in response to that manifest. There are no standard APIs to integrate the presentation decisions made by the platform with a presentation application running in the browser.

However, there is work underway in W3C to add both adaptive streaming and content protection APIs so that a script application will be able to run in any HTML5 browser to perform DRM license acquisition and DASH adaptive streaming under the control of the script application. This will allow the script application to control adaptive heuristics, authorization, load balancing, performance reporting, targeted ad insertion, and interactive presentation and navigation of one or more adaptive presentations. These script APIs will make HTML5/JavaScript DASH applications a viable alternative to Silverlight and Flash across the full range of devices … in the future.

AMWA Releases MXF Commercial Delivery Specification

The Advanced Media Workflow Association (AMWA) has released a new MXF Commercial Delivery specification, AS-12. The constrained version of MXF has been developed to enable more efficient handling of commercials through the many transactional and media processing operations between conception and air.

As broadcasters look to serve commercials to long tail delivery platforms as well as their primary channels, controlling costs is all-important. Many versions may exist of the same commercial, adding confusion to the traffic operations. Versions may be sourced from different distribution routes, arriving with different wrappers and codecs, as well as different aspect ratios.

MXF Commercial Delivery aims to solve two problems: unique identification, and defining a master spot for the creation of long-tail versions.

MXF Commercial Delivery unambiguously identifies the spot through the Ad-ID unique identifier carried in a “digital slate”. Current practice to identify commercials is by the visual slate preceding the commercial. Although a human operator can read this by playing the commercial, it does not lend itself to use by automated systems.

With MXF Commercial Delivery the advertisement identification metadata is carried as a Descriptive Metadata track and serves as a digital slate. The digital slate can be used to reconcile the video and audio components of the commercial with the traffic instruction thus preventing expensive mistakes. The Ad-ID unique identifier ensures that what the advertiser ordered gets to air.

MXF Commercial Delivery, AS-12, is an addition to AS-03, MXF for Delivery. AS-03 defines MXF files optimized for program delivery, and intended for direct playout via a video server. Used together, the specifications allow the agency to supply broadcasters with a master commercial, along with information like closed captions and AFD that the broadcaster can use to create the lower resolution versions appropriate to their long-tail delivery platforms.

The AS-12 metadata or digital slate can be created early in the production process to uniquely identify the commercial. AS-12 carries fields to identify the advertiser, the agency, brand and product, as well as the title. Through the use of the guaranteed unique Ad-ID, the rekeying of identifiers that typically happens today is avoided. House codes used by agencies or broadcasters are replaced with the Ad-ID, avoiding many of the issue of misidentified commercials that have been commonplace.

The Commercial Delivery specification is sponsored by AMWA principal member Ad-ID, a joint venture of the American Association of Advertising Agencies (4A's) and Association of National Advertisers (ANA), and establishes a baseline for improved Operations, Administration, and measurement of advertising assets across the myriad of current and emerging delivery platforms, which when fully deployed, will result in substantial financial gains and improvements in productivity that will flow back to all participants within the supply chain.

Source: Advanced Media Workflow Association

MPEG Transport Stream

The MPEG-2 standard is defined by ISO/IEC 13818 as "the generic coding of moving pictures and associated audio information." It combines lossy video compression and lossy audio compression to comply with bandwidth requirements. The basic structure of all MPEG compression systems is asymmetric because the encoder is always more sophisticated than the decoder.

MPEG encoders are always algorithmic. The better ones are also adaptive, using a feedback path. MPEG decoders are not adaptive and perform a fixed function. This works well for applications like broadcasting, where the number of expensive complex encoders is few and the number of simple inexpensive decoders is enormous.

The MPEG standard provides little information about how encoder processes and operation. Rather, MPEG-2 specifies how a decoder interprets metadata in a bit stream. The metadata tells the decoder the rate the video was encoded, defines the audio coding, and identifies channels and other vital stream information.

A decoder that successfully deciphers MPEG streams is called compliant. The beauty of MPEG is that it allows different encoder designs to evolve simultaneously. Generic low-cost and proprietary high-performance encoders and encoding schemes all work because they are all designed to communicate with the compliant decoder base.

Stream Structures
An MPEG-2 stream can be either an Elementary Stream (ES), a Packetized Elementary Stream (PES) or a Transport Stream (TS). The ES and PES begin with and are stored as files. Individual ESs are essentially endless because the length of an ES is as long as the program itself.

Starting with analog video and audio content, individual ESs are created by applying MPEG-2 compression algorithms to the source content in the MPEG-2 encoder. This process is typically called ingest. The encoder creates an individual compressed ES for each audio and video stream. An optimally functioning encoder will appear transparent when decoded in a set-top box and displayed on a professional video monitor.

A good ES depends on several factors, beginning with the quality of the original source material, and the care used in monitoring and controlling audio and video variables when material is ingested. The better the baseband signal, the better the quality of the digital file. Also influencing ES quality is the encoded stream bit rate and how well the encoder applies its MPEG-2 compression algorithms within the allowable bit rate.

MPEG-2 has two main compression components: intraframe spatial compression and interframe motion compression. Encoders use a variety of techniques, some proprietary, to maintain the maximum allowed bit rate while at the same time allocating bits to both compression components. This balancing act can sometimes be unsuccessful. It is a tradeoff between allocating bits for detail in a single frame and bits to represent frame to frame motion changes.

Researchers are still investigating what constitutes a good picture. Presently, there is no direct correlation between the data in the ES and subjective picture quality. For now, the best way of checking encoding quality is with the human eye, after decoding.

The Packetized ES
Each ES is broken into variable-length packets. The result is a PES containing a header and payload bytes. The header includes information about the encoding process required by the MPEG decoder to decompress the ES.

Each individual ES results in an individual PES. At this point, audio and video information still resides in separate PESs. The PES is primarily a logical construct and is not actually intended to be used for interchange, transport and interoperability. The PES also serves as a common conversion point between TSs and PSs.

Both the TS and PS are formed by packetizing PES files. During the formation of the TS, additional packets, containing tables needed to demultiplex the TS, are inserted. These tables are collectively called PSI and will be addressed in detail later.

Some packets contain timing information for their associated program, called the program clock reference (PCR). The PCR is inserted into one of the optional header fields of the TS packet. Recovery of the PCR allows the decoder to synchronize its clock to the rate of the original encoder clock.

Null packets, containing a dummy payload, may also be inserted to fill the intervals between information-bearing packets.

TS packets are fixed in length at 188 bytes with a minimum 4-byte header and a maximum 184-byte payload. Key fields in the minimum 4-byte header are the sync byte and the Packet ID (PID). The sync byte's function is indicated by its name. It is a long digital word used for defining the beginning of a TS packet.

The PID
The PID is a unique address identifier. Every video and audio stream as well as each PSI table needs a unique PID. The PID value is provisioned in the MPEG multiplexing equipment. Certain PID values are reserved or specified by organizations such as the Digital Video Broadcasting Group (DVB) and the Advanced Television Systems Committee (ATSC) for electronic program guides.

In order to reconstruct a program from all its video, audio and table components, it is necessary to ensure that the PID assignment is done correctly and that there is consistency between PSI table contents and the associated video and audio streams. This is one of the more critical points in a MPEG-2 stream.

There are four other important fields in the TS header. One is the continuity counter. It is a 4-bit field that repeatedly increments zero through 15 for each PID. It’s used to determine if packets are lost or repeated PCR. Second is the discontinuity indicator. It indicates a time base (PCR) and continuity counter discontinuity, which allows the decoder to handle such discontinuities. Third is the random access indicator. It indicates that the next PES packet in the PID stream contains a video-sequence header or the first byte of an audio frame. Fourth is the splice countdown. It indicates the number packets of the same PID number to the splice point when a new PES packet begins.

PSI
During the formation of the TS, additional packets, containing tables needed to demultiplex the TS, are inserted. These tables are collectively called PSI. PSI is part of the TS. PSI is a set of tables required for demultiplexing and sorting out which PIDs belong to which programs.

To identify which audio and video PIDs contain the content of a particular program, a Program Map Table (PMT) must be decoded. Each program requires its own PMT with a unique PID value.

In order to determine which PID contains the desired program's PMT, the Program Allocation Table (PAT) must be decoded. The PAT is the master PSI table with PID value always equal to zero (PID = 0). If the PAT cannot be found and decoded in the TS, then no programs can be found, decompressed, or viewed.

For a set-top box or ATSC tuner to successfully perform the program recovery and decompression process, the PSI tables must be sent periodically and with a fast enough repetition rate that it doesn’t delay channel-surfing viewers. Thus, checking the PSI tables for correct syntax and repetition rate is a vital part of MPEG testing.

Testing PSI involves verifying the accuracy and consistency of PSI contents. As programs change or multiplexer provisioning is modified, some problems may occur. One problem would be unreferenced PID. Packets with a PID value are present in the TS but are not referenced in any table.

If there are no packets with the PID value referred to in a PSI table present in the TS, the problem could be a missing PID.

Another useful PSI test is a check of program content. Just because there are no unreferenced or missing PIDs indicated does not mean that the viewer is receiving the correct program. There may also be a mismatch of the audio content from one program being delivered with the video content from another program. Because MPEG allows multiple audio channels for multiple languages, an air-check can ensure that viewers are receiving the correct language.

It is possible to use a set-top box and television to do the air check, but a better way would be to use an MPEG test set that incorporates all the PSI table checks plus a built-in decompressor with picture and audio display. This would allow you to correlate PSI contents and actual program content as well as allow a quick visual and aural check of ES.

By Ned Soseman, Broadcast Engineering

Production & Exchange Formats for 3DTV Programmes

The purpose of this EBU recommendation is to give technical aid to broadcasters who intend to use current (or future) 2D HDTV infrastructures to produce 3DTV programmes.

Super Hi-Vision at 120fps for Olympics

NHK is developing an 8K image sensor for its Super Hi-Vision Ultra HD system capable of 120 frames per second, which it plans to unveil in Tokyo on 23 April and should be ready for use at the London Olympics.

The new 33-megapixel (7680x4320 pixel) CMOS sensor will use an advanced two-stage (4-bit then 8-bit) cyclic analogue-to-digital converter to deliver a 12-bit image at the higher frame rate (twice that of the current SHV cameras). This architecture also allows it to reduce power consumption, with the ADC drawing 800mW (out of a total drive power of 2.5W). To deal with the high number of pixels involved, the sensor outputs alternating rows of pixels to ADCs on either side, each of which has 48 parallel outputs.

The sensors will probably still get rather warm at these frame rates, so heat management will be a significant issue. Each pixel measures 2.8 x 2.8 microns, and the 26.5 x 21.2mm chip (which is about the width of a Super35mm sensor, but taller) will use a 0.18-micron manufacturing process. The chip is being developed with Shizuoka University, and engineers revealed details of the sensor at the recent IEEE International Solid-State Circuit Conference in San Francisco.

Two SHV cameras will be used to capture parts of the 2012 Olympics, with transmission to three large screens around the UK (plus one in the International Broadcast Centre) and three in Japan. Having the 120fps sensor will reduce motion blur and allow for much better slo-mo replay.

However, as the resolution increases (and 8K is 16 times the resolution of HD), motion defects become much more noticeable. “300 frames per second might just be acceptable, but 600fps would be better,” said colour scientist and camera consultant, Alan Roberts. At 300fps, the material would also be easily compatible with both 50 and 60Hz display.

Using Long GoP compression, which typically combines a group of pictures in half a second, would lead to a GoP of 150 or 300 frames, but this wouldn’t lead to huge bandwidth requirements. “The motion between frames is very small, so the compression is much easier,” he explained, which means the increase in bit rate needed to convey high framerate material is minimal. “The problem is in the shooting and the editing, because of the monstrous data files you have to deal with,” he added.

By David Fox, TVB Europe

The 3D Production Guide

3net, the 24/7 3D network and 3D television production studio, along with joint venture partners Discovery, Sony and IMAX, announced the release of the most complete guide to 3D television production ever assembled.

Featuring stereoscopic expertise from the top producers and technical advisors of the company and its corporate ownership, The 3D Production Guide has now been made freely available to the public via multiple websites.

The 50-page illustrated manual includes detailed information garnered from the combined 50 years of experience in the area of 3D from those who contributed to its creation. The guide outlines in detail all of the facets involved in creating top-quality 3D content for television, from initial workflow planning, to production, post production, stereographic correction and final delivery.

The guide was authored by Bert Collins, Josh Derby, Bruce Dobrin, Don Eklund, Buzz Hays, Jim Houston, George Joblove and Spencer Stephens, with Bert Collins and Josh Derby serving as editors. It will be constantly updated and amended as the dynamics of 3D television production continue to evolve.

Source: 3net

EyeIO: Netflix’s Secret Weapon Against Bandwidth Caps?

Palo Alto, Calif.–based video encoding startup EyeIO left stealth mode on Wednesday with the announcement that it has licensed its technology to one of the biggest players in the online video space. Netflix is using eyeIO’s encoding technology to cut down on the bandwidth of its streams, allowing the company to deliver HD video without busting subscribers’ bandwidth caps or overwhelming networks in emerging markets.

EyeIO has been operating stealthily since the end of 2010, and it was able to win Netflix as a customer last summer. Netflix hasn’t said where and in which capacity it is exactly using the technology it has been licensing from eyeIO, but the company’s VP of Product Development, Greg Peters, said in a press release that eyeIO is “an important part of the technology [Netflix uses] to improve video quality and overcome bandwidth challenges presented by Internet infrastructure.”

Standard-definition Netflix streams can consume up to 2.2 Mbps of bandwidth. Netflix’s 720p HD videos come in at roughly 3.8 Mbps, and 1080p videos go up to 4.8 Mbps. EyeIO CEO Rodolfo Vargas told me during a phone conversation on Tuesday that his company’s encoding technology can achieve better-looking results than most established encoders with 20 percent bandwidth savings and that eyeIO can still deliver similar quality to other encoders with up to 50 percent bandwidth savings. Content in 720p could be streamed using 1.8 Mbps, he explained. The company does this by optimizing the encoding process, which means that the results are regular, albeit smaller, H.264 files that can be played by end users without any need for additional plug-ins.

EyeIO was founded by online video technology veterans; Vargas used to be the senior program manager for video at Microsoft, and one of his co-founders, Robert Hagerty, used to be the chairman and CEO of teleconferencing provider Polycom. The company is privately funded and currently has fewer than 10 full-time employees but is looking to expand over the coming months.

By Janko Roettgers, GigaOM

Large-Sensor Camcorders and DSLRs

When Canon announced its EOS C300 camcorder, which has a 4K2K sensor but does not record 4K2K video, the company also announced it was developing a DSLR that will record 4K2K video. Although it may seem a bit strange that a camcorder designed to employ high-quality cinema lenses is limited to full HD video recording yet a still camera will be able to record 4K2K video, it's not strange given the history of DSLRs.

When large CCD and CMOS chips replaced an SLR's 35mm film, the next logical step was to place an LCD on the digital camera so one could review shots in the field. The next logical step was a “live view” mode that allowed one to view what was being recorded. It was only a small step to compress live view images and record them as video.

Large-sensor digital camcorders have evolved from DSLRs. It is primarily a marketing decision whether to release digital motion picture technology in a still camera package, a camcorder package or both. However, one clear advantage of a camcorder package is space for a mic jack (even XLRs), a headphone jack and manual audio controls.

When a potential buyer who is in the process of learning about 4K2K production and post production encounters the same technology in two different packages, it may prove confusing.

When a videographer shoots with a still camera, he or she will find expected camcorder functions missing. For example, every professional camcorder has some form of ND filtration; DSLRs do not.

A primary differentiator of DSLRs and traditional camcorders is their optical system. This is true for current HD and future 4K2K products.

Frame Size
While video cameras have frame sizes that relate directly to sensor size, such as 2/3in, DSLR frame size relates to 35mm film — in particular, 35mm still film. When shooting 35mm slide or negative film, each 36mm × 24mm image is placed with perforations above and below the frame.

DSLRs with 36mm × 24mm sensors are called full-frame cameras. The Canon EOS-1D X, announced for March 2012, employs an 18-megapixel 28.7mm × 19.1mm sensor. Canon calls it an APS-H sensor.

There are small variations in APS frame size: Canon APS-C (22.2mm × 14.8mm), and Nikon/Sony-C (23.4mm × 15.6mm). Both full-frame and APS sensors, when taking photos, have a 3:2 (1.50:1) aspect ratio. Panasonic uses a slightly smaller sensor for its AF100 camcorder and GH2 still camera called Micro Four Thirds (M43), which has a 1.33:1 aspect ratio and a frame size of 17.3mm × 13mm.



Sensors smaller than a full-frame sensor reduce the potential minimum DOF. Minimum DOF, of course, is a function of the maximum aperture size. A large-sensor camera does not directly provide a shallow DOF.

When a lens designed for a full-frame camera is mounted on a camera with a smaller sensor, the lens' focal length is multiplied by the lens crop factor. (Crop factor equals the ratio of a 35mm frame's 43.3mm diagonal to the diagonal of the image sensor.) A Sony APS-C camera, for example, has a crop factor of 1.5. A 50mm “normal” lens becomes a 75mm tele lens.

When shooting video, a 16:9 window on the sensor is employed. This has three ramifications. First, the viewfinder image will shrink when switching a DSLR to video mode. (This shift can be minimized by shooting 16:9 photos.) Second, the number of pixels read out will be reduced, which is a positive. Third, the lens crop factor will slightly increase. For example, when a Sony APS-C camera is switched to video mode, the crop factor increases to 1.8, thus a 50mm lens acts as a 90mm lens.

The earliest 35mm movie film had a 22mm × 18mm image, with perforations on the sides of each frame. In 1929, the Academy ratio was established. It has a 21mm × 15mm image that has a 1.37:1 aspect ratio. To obtain wide-screen, but not anamorphic, images, a Super 35 frame can be employed.

A 24.9mm × 13.9mm Super 35 frame has a native aspect ratio of 1.79:1 — a perfect match to 1.78 (16:9) HD. It also matches Quad-HD (3840 × 2160 pixels) and almost matches 4K2K, which is 4096 × 2160 pixels — a 1.90:1 aspect ratio. Not surprisingly, frame sizes that come from cinema cameras do not require the use of a 16:9 window when shooting video.


A 24.9mm x 13.9mm Super 35 frame has a native aspect ratio of 1.79:1


Lens Zoom System
While the videographer likely knows that DSLR lenses do not have power zoom, he or she may not know that photo lenses have other issues. For example, the zoom ring may have high friction because of the need to significantly extend the lens when zooming. Pressure exerted to start a zoom while shooting can easily cause a visible disturbance.

Better Sony lenses, such as A-mount lenses that use micro ball bearings, may cause noise that will be picked up by an on-camera mic.

AF System
Photographers are used to trusting auto-focus — even on action shots where the shooter is following a moving subject. When a DSLR's mirror is in the 45-degree-position in order for the shooter to see the subject, AF is possible. A portion of the image passes through a semitransparent area of the mirror, reflects off a small mirror mounted on the back of the mirror and is cast onto a small sensor at the bottom of the camera. The sensor, in conjunction with a processor, sends commands to the lens' AF motor to move to a position calculated to be correct for precise focus. This system is called phase detection AF.


A phase detection AF system uses a series of mirrors, a small sensor at the bottom of the camera
and a processor to calculate precise focus.


DSLRs employ a different AF system when shooting video because the mirror must be continuously up. The processor, therefore, obtains information from the CMOS image sensor, which is why it is called contrast detection AF.

Mirrorless cameras such as the Panasonic GH2 and Sony NEX-5N must use contrast detection AF. (Strictly speaking, digital cameras without a mirror do not have a reflex system and, therefore, are not DSLRs.)

Contrast detection AF systems work by having a microprocessor rapidly command the lens servomotor to step forward and backward by a tiny amount. The processor notes whether contrast increases or decreases. If contrast increases, then current focus is not perfect. Therefore, stepping forward and backward continues. When there is no change in contrast, the current focus is the best possible.

Contrast detection tends to be slower than phase detection and becomes slower at low light levels. And, unless the lens is designed to be quiet, AF noise may be recorded.

Aperture System
Photography lenses are designed to click into key f-stops: f/2.8, f/4, f/5.6, f/8, f/11, f/16 and f/22. Cinema and video lenses are designed so the aperture changes in a continuous manner. One solution is to use camera lenses designed by the camera's manufacturer for video shooting. The other solution is to use cinema lenses.

ND Capability
To obtain a shallow DOF with a large-chip camera under bright light — at the slow shutter speed required for the correct amount of video motion blur — an ND is a must. (ND filtration also will be required to keep the aperture under f/11 to minimize diffraction.) When a camcorder does not have a built-in filter, a shooter has three choices: mount the camera on rails on which a matte box is mounted, attach one of several ND filters to the lens or employ a vario-ND filter.

Lens Mount Type
Both cameras and camcorders that employ large sensors use a lens mount designed to work with their brand of lenses. For example, Sony's NEX family — including the FS100 and VG20 camcorders — uses Sony's E-mount. Sony markets the LA-EA2 adaptor, which enables the use of Sony and Minolta A-mount lenses. The LA-EA2 has a translucent mirror system that provides phase detection AF to many A-mount lenses.

For most interchangeable lens cameras, third-party adaptors are available. These enable you to use your favorite photo lenses on a new camera or camcorder. For example, a Sony NEX camera can use Nikon F, Canon 5D, Leica M, Leica R, Pentax, Konica Minolta MD, Olympus and Contax/Yashica lenses by using an E-mount adaptor.

Only a few adaptors, such as the LA-EA2, provide electrical signals to a lens. Without electrical connections, in-lens optical stabilization, AF and aperture control cannot function and no information from the lens is received by the camera. Therefore, modern photo lenses that send the aperture ring's setting to the AE system cannot do so.

Solutions to these issues include working with still camera and cinema lenses in a fully manual way (which may be a camera operator's first choice) or using a manufacturer's lenses that have electrical contacts.

Bringing it Home
No matter whether you shoot with a still camera or camcorder, images from the sensor must be compressed and recorded. Currently, two codecs are used for recording 4K2K: the Sony F65RAW (16-bit RAW) codec to a docking SRMaster field recorder that records to SRMemory cards or the RED R3D wavelet codec to a REDMAG solid-state drive.

Future 4K2K codec options include H.264 (as a single stream or as four HD streams) and 4K2K versions of current HD formats.

By Steve Mullen, Broadcast Engineering

DPP Unveils Technical & Metadata Standards for File-based Programme Delivery

The Digital Production Partnership (DPP) – a partnership between ITV, Channel 4 and the BBC – has unveiled its new Technical and Metadata Standards for File-based programme delivery in the UK.

Through the DPP, seven major broadcasters (BBC, ITV, C4, Sky, Channel Five, S4C and UKTV), have all agreed the UK’s first common file format, structure and wrapper to enable TV programme delivery by digital file. These new guidelines will complement the common standards already published by the DPP for tape delivery of HD and SD TV programmes.

Working closely with the Advanced Media Workflow Association (AWMA) in the US, the DPP has been the driving force behind the creation of the organisation’s ‘AS-11,’ a new international file format for HD Files. The new DPP guidelines will require files delivered to UK broadcasters to be compliant with a specified subset of this new, internationally recognised standard.

By implementing one set of pan-industry technical standards for the UK, the DPP aims to minimise confusion and expense for programme-makers, and avoid a situation where a number of different file types and specifications proliferate.

The new DPP standards aim to remove any ambiguity during the production and delivery process. A key aspect is the inclusion of editorial and technical metadata, which will ensure a consistent set of information for the processing, review, and scheduling of programmes, as well as their onward archiving, sale and distribution.

As part of the file-based guidelines, the DPP’s member broadcasters have agreed a minimum set of common metadata to be delivered with a file-based programme. And, in a bid to encourage international adoption of its metadata standards, the DPP has worked closely with the European Broadcasting Union (EBU), mapping its minimum set of common metadata to existing ‘EBU-Core’ and ‘TV-Anytime’ metadata sets.

Alongside these new standards, the DPP is currently building a free-to-use, downloadable, metadata application to enable production companies to enter the required editorial and technical metadata easily. The new application is due to launch in spring 2012.

The agreement of these new file based technical standards does not signal an immediate move to file based delivery. Instead, the DPP seeks to provide clarity around digital delivery that will become the expected standard in the future.

During 2012 BBC, ITV and Channel 4 will begin to take delivery of programmes on file on a selective basis. Production companies wishing to deliver by file should discuss this at the point of commission, and seek formal agreement with their broadcaster at the outset of production. After a period of selective piloting, file based delivery will be the preferred delivery format for these Broadcasters by 2014.

Source: Digital Production Partnership

AS-11: MXF for Contribution

AS-11 is a vendor-neutral subset of the MXF file format to use for delivery of finished programming from program producers and distributors to broadcast stations. AS-11 files are intended to be complete and ready for playout.

AS-11 supports playout while the file transfer is in progress, a workflow is referred to as “late delivery”. It is preferable for AS-11 files to be used by playout servers directly without rewrapping of the MXF data structures.

The content may be delivered at the ultimate bit-rate, picture format and aspect ratio, or it may be transcoded at the broadcast station to the required bit-rates and formats. Similar transcoding may be applied to audio and captions; additionally, specific audio and caption tracks may be selected for different broadcast channels.

The content may be pre-packaged for broadcast without further splicing or it may be segmented for ease of insertion or replacement of interstitials.

AS-11 supports SD video encoded as D-10, 50Mbit/s, and HD as AVC-Intra Class 100. Audio can be PCM, AC-3 or Dolby E.

AS-11 defines a minimal core metadata set required in all AS-11 files, a program segmentation metadata scheme, and permits inclusion of custom shim-specific metadata in the MXF file.

Source: Advanced Media Workflow Association

EBU-TT Subtitling Format Published for Industry Comments

The EBU has published a new Subtitling Format specification (EBU Tech 3350). The new format is called EBU Timed Text (EBU-TT) and provides an easy-to-use method to interchange and archive subtitles in XML.

EBU-TT is based on the W3C Timed Text Markup Language (TTML) specification. The EBU format can be seen as a constrained version of the W3C spec, aimed at providing a solution more tailored to broadcast operation. This is especially relevant as broadcasters are increasingly moving to file-based HDTV facilities, where subtitles are created, edited, exchanged and archived together with the content.



The previous EBU subtitling format was EBU STL (Tech 3264), developed at a time when information was still exchanged on floppy discs. However, as many broadcasters still use STL or have archived STL files, great care was taken in the development of EBU-TT to make sure that it provides backwards compatibility with its predecessor.

The EBU is also providing an XML Schema for EBU-TT.

Source: EBU

Tech Upstarts Kicking Glasses in 3D

There's no shortage of innovation from the major TV manufacturers on display at the huge booths at CES: OLED, 4K -- even 8K -- resolution, new interfaces, connectivity, exclusive content. What's absent here, though, are any prototypes to indicate that glasses-free (autostereo) 3D TV is anywhere close to market.

That's not to say that autostereo TV can't be found at CES. It's just coming from smaller companies in smaller booths -- one with barely a booth at all. They are pressing ahead with -- and showing off -- autostereo screens for television, tablets and smartphones while the big makers remain oddly quiet on the topic.

"Consumer electronics companies wanted to get into the home market quickly," said Raja Rajan, chief operating officer of Stream TV, whose booth in Central Hall, of the mammoth Las Vegas Convention Center, is not far from Sony's. "The consumer electronics companies have tremendous financial pressures to get to market with the fastest, easiest technologies."

That is echoed by one of Rajan's competitors, Stephen Blumenthal of 3D Fusion, a late addition to the floor that has one of its models tucked into the 3D Bee booth at the periphery of Central Hall.

"They brought (3D with glasses) to the market as a very straightforward consumer play, and until they burn through the opportunity to make as much revenue off of it as possible, this adventure with the next step is on the back burner," Blumenthal said.

His partner Ilya Sorokin noted, "The 3D with glasses technology was much easier to incorporate into their existing infrastructure because it was already there, and just lying on a shelf."

Both 3D Fusion and Stream TV are using advanced, lens-based tech that, according to Rajan, was abandoned by the big companies.

Rajan said he toured Asia showing Stream TV's screens and its real-time 2D-to-3D converter to major hardware makers, who responded enthusiastically. Stream TV is looking to be a technology provider, not to manufacture under its own name.

"We expect in the next few weeks to start announcing some of the first brands and products rolling out," Rajan said.

He said there is strong interest from Hollywood in the converter box, because it can be built into cable and satellite boxes, enabling all channels to be in 3D. At the same time, Stream's units come with controllers so the consumer can turn the 3D down, or off altogether, for comfort or personal preference.

"Our cost is incrementally 10% to 15% max over the cost of goods for a 2D television," Rajan said. "That's significant because a big re-seller can get into the consumer market at a cost consumers can afford."

MasterImage 3D, which has a solid worldwide business projecting 3D in theaters, is in the South Hall. It has been in the autostereo screen business for some time, and this year is at CES with two screens aimed straight at state-of-the-art mobile devices: a 720p 4.3-inch smartphone display and a WUXGA (1920x1200) display for tablets.

Royston Taylor, exec VP and general manager for MasterImage, said he welcomes the competition from Stream TV, which is also showing tablet screens.

"First, it validates what you're trying to do," Taylor said. "Being on your own is nice in terms of no competition, but it's very lonely in terms of being the only voice saying how great something is. The second thing is competition is always good for the consumer."

Despite strong sales of the Nintendo 3DS, the poor critical response to the 3DS, the HTC Evo 3D phone and the LG Optimus 3D phone have made some makers nervous, Taylor said. He now expects to be making announcements of deals with consumer electronics companies by April and to have gear with MasterImage 3D screens in stores by Thanksgiving.

One hurdle that had to be overcome was the lack of technical standards for judging the quality of a 3D display.

"Right now it's almost entirely subjective," he said. "Big companies won't risk a $250 million phone line on 3D just because it looks nice."

But a French company, Eldim, has come up with a product for testing 3D displays on objective, technical measurements. With standards in place, it will be possible to compare products and establish quality control in manufacturing.

3D Fusion is already selling autostereo TVs for use in digital signage. Blumenthal said the company is selling its turnkey solution, which includes a 42-inch autostereo display, at CES. Cost is $8,000. His sales are to retailers, small mom-and-pop chains, malls. Blumenthal and Sorkin recognize that their company is small and they're in no position to ramp up to consumer volumes on their own. Like Stream TV, they'd be happy to license their technology.

By David S. Cohen, Variety

3-D Cameras for Cellphones

Researchers at Massachusetts Institute of Technology (MIT) have developed a system that uses specially designed algorithms to produce a detailed 3D image with just a cheap photodetector and the processor power found in a smartphone.

Like other sophisticated depth-sensing devices, CoDAC uses the “time of flight” of light particles to gauge depth: A pulse of infrared laser light is fired at a scene, and the camera measures the time it takes the light to return from objects at different distances.

Traditional time-of-flight systems use one of two approaches to build up a “depth map” of a scene. LIDAR (for LIght Detection And Ranging) uses a scanning laser beam that fires a series of pulses, each corresponding to a point in a grid, and separately measures their time of return. But that makes data acquisition slower, and it requires a mechanical system to continually redirect the laser.

The alternative, employed by so-called time-of-flight cameras, is to illuminate the whole scene with laser pulses and use a bank of sensors to register the returned light. But sensors able to distinguish small groups of light particles — photons — are expensive: A typical time-of-flight camera costs thousands of dollars.

The MIT researchers’ system, by contrast, uses only a single light detector — a one-pixel camera. But by using some clever mathematical tricks, it can get away with firing the laser a limited number of times.

The first trick is a common one in the field of compressed sensing: The light emitted by the laser passes through a series of randomly generated patterns of light and dark squares, like irregular checkerboards. Remarkably, this provides enough information that algorithms can reconstruct a two-dimensional visual image from the light intensities measured by a single pixel.

In experiments, the researchers found that the number of laser flashes — and, roughly, the number of checkerboard patterns — that they needed to build an adequate depth map was about 5 percent of the number of pixels in the final image. A LIDAR system, by contrast, would need to send out a separate laser pulse for every pixel.

To add the crucial third dimension to the depth map, the researchers use another technique, called parametric signal processing. Essentially, they assume that all of the surfaces in the scene, however they’re oriented toward the camera, are flat planes. Although that’s not strictly true, the mathematics of light bouncing off flat planes is much simpler than that of light bouncing off curved surfaces. The researchers’ parametric algorithm fits the information about returning light to the flat-plane model that best fits it, creating a very accurate depth map from a minimum of visual information.


Click to watch the video


Indeed, the algorithm lets the researchers get away with relatively crude hardware. Their system measures the time of flight of photons using a cheap photodetector and an ordinary analog-to-digital converter — an off-the-shelf component already found in all cellphones. The sensor takes about 0.7 nanoseconds to register a change to its input.

That’s enough time for light to travel 21 centimeters, Vivek Goyal from MIT’s Research Lab of Electronics, says. “So for an interval of depth of 10 and a half centimeters — I’m dividing by two because light has to go back and forth — all the information is getting blurred together”.

Because of the parametric algorithm, however, the researchers’ system can distinguish objects that are only two millimeters apart in depth. “It doesn’t look like you could possibly get so much information out of this signal when it’s blurred together,” Goyal says.

The researchers’ algorithm is also simple enough to run on the type of processor ordinarily found in a smartphone. To interpret the data provided by the Kinect, by contrast, the Xbox requires the extra processing power of a graphics-processing unit, or GPU, a powerful special-purpose piece of hardware.

“This is a brand-new way of acquiring depth information,” says Yue M. Lu, an assistant professor of electrical engineering at Harvard University. “It’s a very clever way of getting this information.” One obstacle to deployment of the system in a handheld device, Lu speculates, could be the difficulty of emitting light pulses of adequate intensity without draining the battery.

But the light intensity required to get accurate depth readings is proportional to the distance of the objects in the scene, Goyal explains, and the applications most likely to be useful on a portable device — such as gestural interfaces — deal with nearby objects. Moreover, he explains, the researchers’ system makes an initial estimate of objects’ distance and adjusts the intensity of subsequent light pulses accordingly.

Telecoms company Qualcomm has awarded the research team one of $100,000 Innovation Fellowship grants to continue the research.

By Larry Hardesty, Massachusetts Institute of Technology

BetterView

BetterView's up-conversion technology utilizes Super-Resolution (SR) reconstruction that performs a fusion of low quality images into a higher quality result with improved optical resolution. This task encompasses scaling-up of the visual content by introducing true (optical) resolution enhancement.

Several low-resolution images of the same scene. Note they are slightly different from each other.



One HR image (with more pixels, better optical resolution, and less noise), obtained by fusing the previous images.


It has been known for the past 20 years that, in principle, one could take several low-quality images and fuse them into a single, higher-resolution outcome. This has been demonstrated by scientists, adopting various techniques and algorithms. This process became a hot field in image processing, with thousands of academic papers published during the past two decades on the problem and ways to handle it. The classical approach to fuse the low-quality images requires finding an exact correspondence between their pixels, a process known as "motion estimation".

BetterView technology is based on a recently developed and patent-pending novel family of SR algorithms, proposed by a world-leader in this field, Prof. Michael Elad (Technion – Israel Institute of Technology). Elad and his collaborator, Dr. Matan Protter devised the first method that overcomes the requirement for very accurate and explicit motion estimation in previous SR technologies.

Exact motion estimation has been a crucial stage in every earlier SR algorithm, considerably limiting the scenes that can be handled. The new family of SR techniques avoids the exact motion estimation and replaces it by a probabilistic estimate. This enables handling successfully general content scenes containing extremely complex motion patterns. The results are impressive, with no visual artifacts, and the process is completely robust.


Click to watch the video