A site devoted mostly to everything related to Information Technology under the sun - among other things.

Showing posts with label Data Visualization. Show all posts
Showing posts with label Data Visualization. Show all posts

Monday, June 9, 2025

Data as Story - A Proposal

Please find the IP that I developed below on transforming data into stories in 2014 which was disclosed here:

Babak Makkinejad, “Data as Story”, database number 617040, "Research Disclosure", Published in the September 2015 paper journal, Published digitally 21 August 2015 14:06 UT

With the advent of LLMs, my proposal would be easier to implement now.

Subject Matter & Problem

The main idea is this disclosure is the transformation of the relational data – often found in Relational Database Management Systems into something resembling a story; that is: a textual representation of the relational data that is telling a story in a natural language.

This disclosure does not discuss an automated system, rather it describes a semi-automated process in which a human expert would invoke software tools to transform data into a story. 

What is presented in this disclosure is akin to the process of report creation from available relational data using such tools as MS SQL Server Reporting Services, Apache’s Java BIRT etc. in which a human user designs a report – which consumes relational data – using a variety of software tool at the end of which an automated system published that report or makes it otherwise available.

This is not meant as a replacement for other modalities of data presentation such a charts and graphs but is meat as a complementary modality.  However, the presentation of the data as a story will be found by most human beings to be more engaging than looking at reams of form-based data or columnar data as represented by database extracts of MS Excel files.

Once the data is turned into a textual, human language story, it could be read out to a human being mechanically, or it could be automatically translated to a different human language. 

Solution

This solution crucially and fundamentally relies on the prior art embodied in the software tool called Dramatica (www.dramatica.org) and its Theory of Story (http://dramatica.com/theory/book).  The Theory of Story is briefly sketched out below.

Introduction to Dramatica’s Theory of Story

The Theory of Story embodied in Dramatica models a story as a single at work finding a solution to a single problem.  This is very analogous to the situation in Business Intelligence arena when different users, in trying to answer different questions, ask for different reports out a database system.  The BI users, in other words, are trying to solve a problem.

We have 4 main areas:

  1. The overall story
  2. The main character through whom we see everything
  3. Impact character
  4. The dynamics of the Impact Character vs. Main Character       

In each of the above areas one has to answer these essential questions:

  1. 1.    Main character’s resolve: will he change or remain steadfast (no story if none of the characters in the story change)?
  2. 2.    What drivers the story –actions or decisions
  3. 3.    What is the main problem class of the story – Fixed Attitude, Manipulation, Situational, or Activities?
  4. 4.    What is the main concern of the story – Past, Present, Future, and Dynamic (How things are changing)?
  5. 5.    What is the overall story issue – Openness, Delay, Choice, Pre-conceptions?
  6. 6.    What is the overall story problem – Control, Help, Hinder, Uncontrolled

Additionally, there could be multiple development lines in each story each with their own thematic arguments.  Themes are perspectives and thus could represent data as viewed from different view point of other story characters.

An argument’s topic may be further explored through dialogue, images, charts, pictures music etc. that complement the story. 

These later supporting material such as charts and graphs ties us to the common data representations via the applications of statistical algorithms and standard charting techniques.

Dramatica’s Theory of Story posits the existence of an overall story symptom and an overall story response; each story consists of a Problem, a Direction, a Focus, and a Solution.  The Problem is finally recognized some time near the climax of the story.  “Success” means the problem is replaced with a “Solution”.  “Failure” means that the problem is persisting.

Drmatica’s Theory of Story further posits that each story could contain up to 8 archetypes:

  • Protagonist vs. Antagonist
  • Guardian vs. Contagonist
  • Reason vs. Emotion
  • Side Kick vs. Skeptic

The Dramatica Structural Matrix

This is a framework for holding dramatic topics pertinent to Genre, Plot, Theme, and Character in relationships that describe their effect upon one another.  There are 4 Classes, Universe, Physics, Psychology, and Mind.  Each class contains 4 Types, and 16 variations (4 each) for each Type.  Each of those 16 Variations, in turn, contain 4 Elements for a total of 64 elements.

During the process of story-forming, these topics (called "themantics") are re-arranged much as a Rubik's cube might be scrambled, all in response to the author's choices regarding the impact they wish to have on their audience. As a story unfolds, the matrix unwinds, scene by scene and act by act until all dramatic potentials, both large and small have been completely explored and have fully interacted.

It is during this phase of story-forming that the relational data – based on their semantics (i.e. the meaning of the data columns in the database) are mapped into these 64-elements for each of the 4 Dramatica Classes.

Approach to Story Construction

Enterprises, commercial, governmental or non-profit, internally execute a set of (business) processes.  This is where the work for the story creation starts.  Examples of such processes are Human Resources, In-patient Management, Out-patient management, Manufacturing Quality Management and very many more.

One selects an existing business processes which is being executed and for which one wishes to tell a story.  That is, a specific business problem or question would be addressed via the story that is being developed.

For this process – or indeed any process – then tries to find the answer to the following questions:

 

1.       How

2.       What

3.       When

4.       Where (to/from)

5.       Who

6.       Whom

7.       Whose

8.       Why

 

Not all of these questions could potentially have answers within an arbitrary business process but some of them will have answers by necessity.

For example, for a Human Resources Management process, the questions could be:

What: role, title

When: hired, left the company, promoted, demoted, reprimanded, recognized, rewarded

Which: salary, rewards, taxes, expenses

Who: The Specific Employee (the Protagonist)

Where: Head-Quarters, Working-from-Home, Branch Office

This step may be automated via software Wizards that guide the user in determining the answers.  Such an automated systems will consume the relational data that supports that business process.  This identification may be based on automated inference or via data dictionaries available for the targeted process.

[A data dictionary contains the semantics of the data elements in the database; it may be viewed as an Ontology for that process – or it could be a subset of the larger Ontology of the Entire Enterprise.]

Next, with the answers to the above questions, and in conjunction with the data dictionary for the database tables, the dominant Class of the story may be selected.  It might be that the story to be developed does not have a dominant Class and all Classes need to be included to present the data.  An automated “Semantic Extraction” tool may be used to facilitate the assignment of the data fields in the database tables to these 4 Classes, their 4 Types, 16 Variations, and 64 Elements.  Alternatively this step could be performed manually. 

This is the step that ties the RDBMS data to Dramatica’s structures.

In practice, the meta-data from the data dictionary may not be sufficiently numerous to cover all 256 bins (Elements) that are available for all the 4 Classes.  Or, alternatively, there could be multiple meta-data elements (concepts) that are mapped to the same Element.  It is a judgment call by the story-teller, looking at the requirements of the story, to decide which meta-data elements to keep and which ones to discard.

The story, ultimately, is a report and must supply answers to the business questions/problems that are posed by its consumers/users.

At this point, the storyteller is in position to utilize a system based on Dramatica and its Wizards, to guide him through the construction of the story. 

For example, the story might be one that is required to tell what happened to a (heart) patient admitted to the emergency room.  Within the Theory of Story of Dramatica, the storyteller would identify the patient (at this stage an unknown person) as the Protagonist, (Heart) Disease as the Antagonist, the (Heart) Surgeon as the Guardian, and Pre-existing Medical Conditions as the Contagionist.  The story would tell the how/when/where/who was admitted, the initial diagnosis, the climax (open heart surgery) and the recovery and discharge as well as other details per the requirements (members of the surgical staff, length of the operation, type of procedures, etc.).

Another example could be Manufacturing Quality Management process in which a type of widget is identified as the Protagonist, Manufacturing Process is considered as the Antagonist, Quality Engineer is identified as the Guardian and the Contagionist could be a production worker or the production machinery.

It should be noted that the stories that are being discussed in this disclosure are all generic, the identity of the patient, or the disease or the widget type are left undefined.  In the sense of RDBMS Reports, these stories may be understood as parameterized reports.

Dramatica already has Wizards that guide one in the construction of one’s story.  However, for this disclosure, it is envisioned that Dramatica’s Wizards and perhaps engine would be augmented in such a manner as to facilitate the construction of stories in which the 8 archetypes are not necessarily human being but could be things or processes.

For example, SQL queries to automatically pull data for an unknown patient, with an unknown disease could be formulated via Wizards during the story-forming stage.  The (augmented) Dramatica will then generate the story per the relational data, its meta-data, and story-teller’s decisions.

The user could then access an online system and request a story that tells what happened to Mr. Smith, with heart disease, who entered the emergency room of the General Hospital between the hours of 8:00 PM to 8:00 AM during the month of August – if any.  The system will generate a textual story as discussed in the above disclosure which could be printed, emailed, turned into an audio-file or exported to a suitable format for printing such as MS Word or Adobe PDF.

Possible Modifications

Extension to non-SQL, unstructured data.

Extension to telling multiple stories at once – by following different business processes within the same story.  For example, while telling the story of the treatment of a heart disease patient which is within the In-Patient Management process, one can also tell the story of specific surgeon who operated on him within the Human Resource Management process.

A story may be incorporated into a more customary report which contains tabular data and charts.  Alternatively, the story itself, in its published form when its parameters are specified, may contain such tabular data and charts.

The Dramatica Documentation

 

  

Saturday, February 18, 2017

Recent Advances in Data Science

Last year, Zoubin Ghahramani published an algorithm that automates the job of a data scientist, from looking at raw data all the way to writing a paper.

His software, called Automatic Statistician, spots trends and anomalies in data sets and presents its conclusion, including a detailed explanation of its reasoning.

The paper is @ http://www.nature.com/nature/journal/v521/n7553/full/nature14541.html
 

I think the implications are quite clear.

Already, there is: https://www.automaticstatistician.com/about/#

Tuesday, July 26, 2016

Data Visualization and Art

My denied patent application for data visualization using computer-animated figure movements:


http://www.google.com/patents/US20100039434


and my patent for data visualization using data painting of an asymmetrical facial image  in the style of Pablo Picasso:


https://www.google.com/patents/US20090110244?dq=Data+Visualization+asymmetrical+faces&cl=en

Monday, October 26, 2015

Data Visualization with Dance

This is an idea that I tried to get patented but was denied by USPTO  (see please http://www.google.com/patents/US20100039434 ).  Nevertheless, I think the idea has merit (say in visualization of the dance of life within a cell) and I am sharing it here, hoping others would pick it up and run with it.


Data Visualization with Dance


The core idea is to visualize statistical properties of single and multi-variate data by means of dance movements of computer generated animation figures.
Traditional visualizations utilize basic dimensions of graphical representation to portray multi-modal time-varying data: color, shape, size, and location. The dimensionality of these representations is limited and their aesthetic quality on the whole is generic. Moreover, the visualization is often static, the time dimension is "frozen-out".  This IDF aims to make improvements to the visual display of time-varying data.
The first computerized dance notation system, which displayed an animated figure on the screen which performed the dance moves specified by the choreographer, was the DOM dance notation system, created by Eddie Dombrower on the Apple II personal computer in 1982. (See Dance Notation Journal, Fall, 1986, 4(2) pp. 47-48.)  This solution consisted of a single figure, and the movements of that figure were not based on any external data sources or their statistical properties.
There are numerous COTS packages used in Web Design and Electronic Games industry that enable one to create figures and to animate and display those figures.  I envision leveraging such tools in building this system.
I propose there to use the language of dance in particular and the language of movement in general to visualize statistical properties of single and multi-variate time varying data.  These streams of data may originate from sensors, from real-time structured and un-structured data, or may be the outputs of forecasting calculations.
My approach consists of several steps.  Through a user interface, a user selects the data stream(s) that he wishes to visualize using dance. Next, he uses software based tools to create a figure the movements of whose limbs are going to indicate the data and its statistical properties.  For each data stream he will create a distinct figure.  The figures will differ in color, shape, and size thus visually indicating the distinct time-varying data streams that are being visualized.
The system saves the results of the figure creation.
Next, the user, will choreograph these figures based on statistical properties of the data.  For example, the user may decide to assign specific dance moves to those figures for which the data streams are beyond a certain threshold as defined by the mean-value of the data, or as defined by degrees from the standard deviation of that mean.  Or the user may decide to indicate those data streams that have positive correlations with one another by figures that dance with and around each other in close proximity of one another- depending on the degree of the correlation.  In a similar manner, the user may choreograph these figures with dance steps so that other statistical properties of data are thereby indicated - higher moments of the distribution function and so on.
The user will input the choreography in the form of one of the available dance notation systems such as Labanotation system or the Sutton Movement system.  An embodiment of this invention using the Sutton Movement system enables inclusion of skate-boarding, pantomime, gymnastics and other such activities as ways of indicating time-varying data.  The user connects his dance choreography with the statistical properties of the time-varying data through this interface.
The system saves the results of the choreography which consists of dance movements as well as statistical properties that trigger those movements and guide them.
At this stage, the user has accomplished 3 tasks: he has created figures, he has identified his data streams with specific figures, and he has choreographed dance moves for each figure based on the statistical properties of (potentially all of) these time-varying data streams.
Next the user indicates to the system to animate these figures based on the input data streams and the dance moves that were choreographed and saved earlier.  The system will begin processing the data stream(s) and compute the statistical properties of the time-varying data.  Based on these statistical properties, the system will automatically load the figure from its data store and invoke the corresponding dance movements for each figure (data stream).  The system will then display animated dancing figures on a display device that indicate the statistical properties of the data stream(s).
The display device and the delivery of the images is not part of this; those task scan be accomplished through WEB, Client-Server, Mobile, or other architectures and technologies.


Web References:


Journal Articles:


  1. Eddie Dombrower, Dance Notation Journal, Fall, 1986, 4(2) pp. 47-48.
  2. M. Cunningham, Changes/Notes on Choreography, F. Starr, ed., Something Else Press, 1968.
  3. D. Tolani, A. Goswami, and N. Badler, "Real-Time Inverse Kinematics Techniques for Anthropomorphic Limbs," Graphical Models, vol. 62, no. 5, 2000, pp. 353-388; .
  4. M. Van de Panne, "From Footprints to Animation," Computer Graphics Forum, vol. 16, no. 4, 1997, pp. 211-223.
  5.  L. Wilke et al., "Animating the Dance Archives," Proc. 4th Int'l Symp. Virtual Reality, Archaeology and Intelligent Cultural Heritage (VAST), Eurographics Assoc., 2003, pp. 91-99.
  6. M. Nakamura, "Text Representation of Labanotation Data for Computer Based Motion Analysis," presented at the World Dance Assoc./Int'l Council of Kinetography Laban/Congress on Research in Dance Int'l Conf., 2004; (http://www.imb.is.ritsumei.ac.jp/~hachihachi_e.html)
  7. A. Hutchinson Guest, Labanotation: The System of Analyzing and Recording Movement, Taylor and Francis, 1987.
  8. Tom Calvert , Lars Wilke, Rhonda Ryman, Ilene Fox, “Applications of Computers to Dance”, IEEE Computer Graphics and Applications  March/April 2005 (Vol. 25, No. 2)   pp. 6-12.
  9. “A Prototype Program for Visualization of Dance Performances Using 3DCG Motion Database”,
  10. Umino Bin(Fac. of Socil., Toyo Univ.)   Soga Asako (Ryukoku Univ. Fac. Sci. and Technol.)  
  11. IPSJ SIG Technical Reports, ISSN:0919-6072, VOL.2006;NO.85(CH-71);PAGE.41-46(2006)
 

Books:


Patent:


“Dance visualization of music”, United States Patent 6717042

Abstract:

An apparatus is equipped to provide dance visualization of a stream of music. The apparatus is equipped with a sampler to generate characteristic data for a plurality of samples of a received stream of music, and an analyzer to determine a music type for the stream of music using the generated characteristic data. The apparatus is further provided with a player to manifest a plurality of dance movements for the stream of music in accordance with the determined music type of the stream of music.

 

Thursday, May 7, 2015

Automatic Statistician


 
Reports are generated automatically and have natural language describing patterns in data.

Thursday, July 28, 2011

Visual Display of Quantitative Data

US Debt (excludes the 80 trillion dollars or so of un-secured financial instruments such as derivatives.)

http://www.wtfnoway.com/

Friday, January 7, 2011

200 countries over 200 years in 4 minutes.

A very nice data visualization exercise of transition of 200 countries over 200 years in just 4 minutes.

http://www.flixxy.com/200-countries-200-years-4-minutes.htm

Sunday, March 7, 2010

Pivot for Web Exploration

An interesting presentation and web data visualization tool by Gary Flake (a Microsoft Technical Fellow):

http://www.ted.com/talks/gary_flake_is_pivot_a_turning_point_for_web_exploration.html

The tool is available from Microsoft Labs.

Saturday, June 13, 2009

HIPerWall

HIPerWall is a wall built of numerous high-definition monitors, each with its own imbedded computer for displaying standard and large (up to one gigabyte or larger) graphic images, high-definition (HD) digital movies, and streaming content from video cameras and other live feeds.

Built at Calit2 (California Institute for Telecommunications and Information Technology) at UC Irvine, the Highly Interactive Parallelized Display Wall (hence HIPerWall) has the ability to display 200 Megapixel images.

It is designed to visualize enormous data sets and allows viewers to see detail, with 100 dots per inch on the screens, while retaining the context of an overview by seeing surrounding data (also in high detail). This allows a group of scientists to collaborate, share detailed information, while still keeping the big picture.


Learn more @ http://hiperwall.calit2.uci.edu/

Saturday, October 18, 2008

3D Printers & Data Visualization

The availability and advancements in three-dimensional (3D) printing technologies makes it possible to add additional dimensions for the visual displays of the data. The 3D technologies can create visual (3D) images which have color, height/depth attributes as well as texture attribute. Hence data may be visualized and communicated with sight & touch.

I suppose future enhancement could be made in which the 3D printed image would also use sound & music – in combination with sight & touch – to convey other properties of data. [This is not as far-fetched as you might think; there are already 3D printers that produce working mechanical mechanisms.]

3D Printers

An overview of 3D printers is given at:

http://en.wikipedia.org/wiki/3D_printing

One may think of a 3D printer as a 3D inkjet printer that deposits droplets of plastic, layer by layer, gradually building up an object of any shape. Fabbers have been around for two decades, but they've always been the pricey playthings of high-tech labs — and could only use a single material. A Fab at Home kit costs around $2400 and allows users to print anything from Hors d'Oeuvres to flashlights."


3D Visualization

This document below discusses the application of 3D printers to cartographic data visualization

http://www.bbr.bund.de/nn_103116/DE/Raumbeobachtung/Werkzeuge/Visualisierung/Veroeffentlichungen__Artikel/visualizationsurfaces,templateId=raw,property=publicationFile.pdf/visualizationsurfaces.pdf


Example of 3D Printing Technology For Data Display

This is an example of a 3D printing product and its application to the display of geo-spatial data. It is from http://www.directionsmag.com/article.php?article_id=2034&trv=1

Friday, October 10, 2008

Viz Designer

SPSS Viz Designer creates dynamic visualization based on the most appropriate chart or graph for that data set. Find it @ http://www.spss.com/VizDesigner/

The product is based on the "Grammar of Graphics" by Leland Wilkinson.

Saturday, September 20, 2008

Starvation .Net

Check out www.starvation.net. I am impressed by the creative way that the world map in that site conveys grim quantitative information.

About Me

My photo
I had been a senior software developer working for HP and GM. I am interested in intelligent and scientific computing. I am passionate about computers as enablers for human imagination. The contents of this site are not in any way, shape, or form endorsed, approved, or otherwise authorized by HP, its subsidiaries, or its officers and shareholders.

Blog Archive