Monday, February 18, 2019

Some staffs are more equal than others

After talking the matter over with Barbara, I think the problem we are having with phrase agreement is actually quite interesting. It boils down to differences in how Justin and Anne are marking the lower staff. Right hand agreement is still quite solid.

Here are the summary numbers for the first batch of revisions Justin gave me compared to Anne's (over sections 2.1, 3.1, 4.1, 5.1, and 6.1). "Staff 1" is the upper staff. "Staff 2" is the lower. We are looking at the "kappa" score. A kappa score of 1.0 reflects perfect agreement. Anything over 0.8 is fantastic. No one is going to argue with agreement like that. But many people (like me and Barbara) will be suspect of anything below 0.67.

Staff 1 kappa: 0.878523692967741
Agreement: 1255/1262 One: 31 Other: 28
Staff 2 kappa: 0.409898477157359
Agreement: 914/945 One: 33 Other: 21
Overall kappa: 0.655041358347147
Agreement: 2169/2207 One: 64 Other: 49

Anne is "One," and Justin is "Other." These are the complete phrase counts annotated by each of you. Note that Justin has 12 fewer phrases marked on the lower staff than Anne, but only 3 fewer on the upper staff.

It strikes me that all of this may be trying to tell us that phrasing in these left-hand (accompanying) voices are inherently more ambiguous than in the conventionally "melodic" voices--or that two distinct principles are being applied here. It seems quite plausible that a bass line may have its own agenda, over which an independent melody can form its own thoughts. Justin seems to see the situation more like that, while Anne seems to have a stronger sense that accompanying lines reflect the phrasing of the lines they are accompanying.

I am not saying either of you is right or wrong, but we need to see if we can find a way to get you to agree. I am also curious to see if there is any discussion in the (music theory) literature of the interplay between melodic and accompanying phrases.

Also, we did not see this disagreement in Sonatina 1. Why would that be?

I attach a zip file with the phrasing each of you originally submitted (except for 1.1 and 1.2, which have Anne's corrections). Please follow the "Comparing Annotations" procedure at the bottom of this blog post to review the data in abcDE.

I have removed the sub-phrase and motive marks, so we don't have to struggle differentiating those little vertical lines. I would be interested in hearing Anne and Justin offer a rationale for their own and maybe each other's segmentation, with reference to specific examples.

We can of course sidestep this issue for the time being by ignoring the lower staff, but I am not quite ready to do that yet.

Monday, January 21, 2019

Clementi corpus building

The goal is to mark phrases in the Clementi Op. 36 sonatinas, as in this PDF file.

Please use this web interface (abcDE) to enter the phrase marks. Click the cog button on the right and fill in the fields. You are the authority and the transcriber. After you set yourself up, you will need to reload the page in your browser.

I have divided the sonatinas into 27 sections. (Anything between double bar lines got its own section.)

Here are all the section files to be annotated:
11.abc 21.abc 31.abc 35.abc 44.abc 53.abc 62.abc
12.abc 22.abc 32.abc 41.abc 45.abc 54.abc 63.abc
13.abc 23.abc 33.abc 42.abc 51.abc 55.abc 64.abc
14.abc 24.abc 34.abc 43.abc 52.abc 61.abc
(The first number is the sonatina number. The second is the section number. So 24.abc has the fourth section of the second sonatina.)

To open a section file, click the globe button. It should show you a URL like this:
https://dvdrndlph.github.io/didactyl/corpora/clementi/11.abc
To select a different file, change the number. Click OK.

To mark phrases, select the last note included in the phrase, and type a period ("."). This will display a vertical bar after the note to mark the phrase.

The browser should remember your work, but to be safe (and to share your work with me), you will click the eyeball button, copy all of the displayed text, and save it in a text file. Or just paste the text in an email and send it to me. That might be easiest.

You should also mark sub-phrases (";") and motives (",").

I am primarily interested in the complete phrase markings right now, so you can ignore the other marks. But it wouldn't hurt to add the more granular markings while you are in the neighborhood. Of course, we are actually trying to identify how a piece is chunked with respect to fingering, not musicality. But one study suggests these are the same thing.

Regarding annotation guidelines, this is what we have:
The primary task is to demarcate phrases in the score. Mark the notes that end complete musical thoughts, typically supported by the presence of a cadence.

Each voice in a piano score may have its own independent phrasing. However, when a lower voice is accompanying an upper voice, the lower will typically end a phrase around the same time as the upper, coordinating to create the sense of cadence. Pay special attention when deviating from this general rule.

You should also mark sub-phrases and motives. Motives are the smallest definable ideas in music. Sub-phrases are more developed but remain incomplete, perhaps like clauses in a sentence.

Note that the phrase mark subsumes the sub-phrase and motive marks, and the sub-phrase mark subsumes the motive mark. That is, a sub-phrase boundary implicitly terminates a motive, and a phrase boundary implicitly terminates a sub-phrase and a motive. So at most one mark should follow any given note.

Differentiating sub-phrases and phrases can be difficult. When you have serious doubts about this distinction, prefer short phrases over long ones.

Always annotate the last phrases in the piece on both staffs. If a piece ends in the middle of a phrase on either staff, we need to change how the larger work has been divided into sections. [Let Dave know.]

Comparing Annotations

I may email you a zip file, so you can compare each other's annotations. Please follow these steps:
  1. Unzip the contents.
  2. Open abcDE.
  3. Click the cog button.
  4. Set Restore Data to Never.
  5. Click Close.
  6. Refresh your browser window.
  7. Open the individual files (through the "Choose File" button).
To see Anne's annotations, select 1 from the Sequence spinner. To see Justin's annotations, select 2.

When you are done, do the following or abcDE will never remember fingerings that you have entered previously in the browser:
  1. Click the cog button.
  2. Set Restore Data to Always.
  3. Click Close.
  4. Refresh your browser window.
Note that relying on the browser to remember your work is fraught with peril. Whenever you have work you want to save for later, you need to do this:

  1. Make sure the Annotated checkbox is not checked.
  2. Click the eyeball button.
  3. Select all of the text.
  4. Copy it to a text file using TextEdit (Mac) or NotePad (Windows).
  5. Save that file.
This is awkward, but I don't have an easier way to do this at the moment.

Thursday, January 10, 2019

Clustering of editorial advice

A possible enhancement

Again, say we have the following fingerings in our gold standard:
  1. 1234 131 (1 annotator vote) .1
  2. 1xx4 1x1 (3 votes) .3
  3. 12xx x3x (2 votes) .2
  4. 1x1x 1xx (4 votes) .4
How much credit should a model get for suggesting 1234 131? Or 1214 131? That is, how likely is it that the user will accept the advice?

Do we just sum over all the matches? So 1234 131 would get .1 + .3 + .2 = .6? And 1214 131 would get .3 + .2 + .4 = .9? But don't we have more evidence that 1234 131 is a good fingering?

We proposed amplifying fingering sequences based on how many actual annotations it has:
  1. 1234 131 (7 notes x 1 annotator = 7 votes) .189
  2. 1xx4 1x1 (4 x 3 = 12 votes) .324
  3. 12xx x3x (3 x 2 = 6 votes) .162
  4. 1x1x 1xx (3 x 4 = 12 votes) .324
Or should we reduce the contribution of a sequence if it is shared with multiple sequence sets ("clusters")? Otherwise, we are over-representing the signal from less discriminating voices and in general over-stating the likelihood that a user will be satisfied.

So 1234 131 would get .189 + .324/2 + .162/2 = .189 + .162 + .081 = .432. And 
1214 131 would have .162 + .081 + .324 = .567.

My confidence in someone who seems to agree with me is diminished every time I see this person agree with someone I disagree with. So maybe we divide the amplified weight by 1 plus the number of wrongsters we detect agreeing with an editor we see agreeing with the system suggestion.

So 1234 131 would get .189 + .324/2 + .162/2 = .189 + .162 + .081 = .432. And 
1214 131 would have .324/2 + .162/2 + .324 = .567. 

Or actually we should consider the amplified weight of the offending fingerings.

Conversely(?), what if I agree with someone and then I see this person disagree with someone I also agree with? Is agreement transitive. Does A=B and B=C imply A=C? This is not possible, unless I need more coffee. The system suggestion never includes wildcards, and the only way to disagree is by failing to match a non-wildcard element. If I (the system) agree with two people, the two people must agree with each other.





Incomplete editorial fingerings

Say we have the following fingerings in our gold standard:
  1. 1234 131 (1 annotator vote) .1
  2. 1xx4 1x1 (3 votes) .3
  3. 12xx x3x (2 votes) .2
  4. 1x1x 1xx (4 votes) .4
How much credit should a model get for suggesting 1234 131? Or 1214 131? That is, how likely is it that the user will accept the advice? How confident are we that the advice is likely to be good?

Do we just sum over all the matches? So 1234 131 would get .1 + .3 + .2 = .6? And 1214 131 would get the same: .2 + .4 = .6? But don't we have more evidence that 1234 131 is a good fingering?

We could amplify a fingering sequences based on how many actual annotations it has:
  1. 1234 131 (7 notes x 1 annotator = 7 votes) .189
  2. 1xx4 1x1 (4 x 3 = 12 votes) .324
  3. 12xx x3x (3 x 2 = 6 votes) .162
  4. 1x1x 1xx (3 x 4 = 12 votes) .324
Or I am thinking the more specific sequences should be amplified by less specific sequences that do not contradict them:
  1. 1234 131 (1 + 3 + 2 = 6 votes) .353
  2. 1xx4 1x1 (3 + 2 = 5 votes) .313
  3. 12xx x3x (2 votes) .125
  4. 1x1x 1xx (4 votes) .25
Or a combination, like this:
  1. 1234 131 (7 + 12 + 6 = 25 votes) .4
  2. 1xx4 1x1 (12 + 6 = 18 votes) .3
  3. 12xx x3x (6 votes) .1
  4. 1x1x 1xx (12 votes) .2
Shouldn't #3 and #4 reinforce each other somehow?

Also, shouldn't the amplification run both ways? If a complete annotation agrees with a partial, what does that tell us? An editor, who is pitching his advice to a general audience in what we assume is a minimally idiosyncratic way, says these milestone fingerings are the most important.

Given this, I am leaning back toward the simple summing approach I mentioned first, but amplifying by digit. So 1234 131 would get .189 + .324 + .162 = .675, and 1214 131 would get .162 + .324 = .486. This at least passes the smell test.

It strikes me that the sparseness of editorial scores may actually be a blessing. All of my trevails with edit distances seem moot in editorial scores, or at least rendered less pertinent. The editor implicitly tells us which specific annotations are most important and which are free to vary. Using Hamming distance here is less controversial: a digit either matches or it doesn't. This is clearly justified if we assume the editor is being explicit in the advice that actually appears.

Are we really not striving to model the behavior of a good editor and less that of a good pianist?

Tuesday, January 8, 2019

The semantics of editorial fingering

Dear Anne and Justin:

Fingering annotations are typically sparse in editorial scores. This is the case even in pedagogical works intended for beginners. Editors understandably do not want to clutter their scores with unnecessary information. However, this state of affairs makes it difficult to leverage editorial scores as sources of fingering data suitable for training and evaluating computational models.

This raises a number of questions for me.

What is the intent of the editor? Is it to provide complete guidance in a compact format, as appears to be the case in beginner scores? (I remember being puzzled and a little irritated by the missing annotations. Why did I have to interpolate when I don't know what I am doing?!) Or is it to convey major transitions only (hand repositionings) and leave other "minor" decisions to the performer? Or is it to provide advice only in areas of special difficulty and to leave the rest to the performer's discretion? Or is it a combination of these intents, with the emphasis varying over the length of a piece?

How much do you think two pianists would agree when transforming a typical sparsely annotated score into a completely annotated score? Would this vary by editorial intent? (We have data we could use to tease out some answers here, I think. In the WTC corpus, we should have complete human data overlapping sparse editorial data that agree on the notes marked by the editor.)

Are there rules you apply to "fill in the blanks" in editorial scores? Are such rules discussed or codified in the pedagogical literature? Would doing so constitute a potential contribution to the pedagogical literature? Is this something you would like to pursue on its own merits?

Some time ago, I tried to brainstorm on this topic. ("MDC"--Manual Data Collector--is what I originally called the abcDE editor.)

Filling in the blanks would definitely be in order to augment our training data for machine learning. But the existence of blanks is also cramping my style in validating my latest novel evaluation metric. (It involves clustering fingering advice according to how "close" the individual fingering suggestions are to each other. This idea of closeness, already somewhat controversial, is even more strained when the suggestions are riddled with blanks.)

Thursday, July 19, 2018

2018-07-19 status

Done

Administrivia

  1. Booked trip to ISMIR 2018 in Paris.

Model Building

  1. Finailized "Corrected Parncutt" implementation for everything by cyclic patterns.
  2. Confirmed inconsistencies in published Parncutt results:
    • Small-Span and Large-Span penalties are conflated.
    • Small-Span penalty definition is inconsistent.
    • Position-Change-Count and Position-Change-Size penalties are incorrect.
    • Penalty totals are incorrect.
    • Explanatory example has confusing/incorrect costs.
  3. Completed full regression test of code base.
  4. Met with Alex Demos and agreed to co-author paper for Music Scientiae on Parncutt, Corrected Parncutt, Improved Parncutt, and how to tell them apart. Will dry run some of this material in late-breaking paper at ISMIR.

    Doing

    1. Implementing support for, and clarifying definition of, cyclic pattern constraint in Parncutt. (Should also do this for Sayegh and Hart.)
    2. Double-checking pruning mechanism in Parncutt.
    3. Writing up findings on "Corrected Parncutt" model, initially for ISMIR submission.
    4. Adding mechanism to learn weights for "Improved Parncutt" rules from training data.

    Struggling

    1. How does one compare two ranked lists of sequences to a third and claim one of the two is more similar to the third in a statistically significant way? That is the big question.

    In Scope

    1. Implementing crude automatic segmenter.
    2. Developing staccato/legato classifier.
    3. Demonstrating improved Parncutt via #1 and #2.
    4. Debugging Sayegh model, which produces results inconsistent with training data.
    5. Developing better test cases for Sayegh.
    6. Updating abcDE to support manual segmentation.
    7. Completing and polishing abcD for entire Beringer corpus.
    8. Defining initial benchmark corpora and evaluation methodology.
    9. Implementing convenience methods for reporting benchmark results.
    10. Moving Beringer corpus to MySQL database.
    11. Enhancing Parncutt, following published techniques and pushing beyond them.
    12. Enhancing Hart and Sayegh to return top n solutions.
    13. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
    14. Adding support to abcDE for annotating phrase segmentation.
    15. Debugging Dactylize 88-key circuit.
    16. Collecting fingering data from JB performances in Elizabethtown.
    17. Completing Dactylize II circuit.
    18. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
    19. Defining procedure for sanity test of production automatic data collector (including Beringer data).
    20. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
    21. Implementing end-to-end machine learning experiment, using Beringer abcD data.
    22. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

    Friday, June 22, 2018

    2018-06-22 status

    Done

    Model Building

    1. Re-implemented Parncutt cost functions for both hands.
    2. Identified possible inconsistencies in published Parncutt model description.
    3. Added feature to track more granular cost details in Dactyler models.
    4. Tracked individual rule costs to facilitate analysis of Parncutt results.
    5. Completed successful regression test for hacked-up Didactyl code. The APIs they are a-changin'.

      Doing

      1. Testing Parncutt cost functions.
      2. Reproducing original Parncutt results.
      3. Adding mechanism to learn weights for Parncutt rules from training data.

      In Scope

      1. Implementing crude automatic segmenter.
      2. Developing staccato/legato classifier.
      3. Demonstrating improved Parncutt via #1 and #2.
      4. Debugging Sayegh model, which produces results inconsistent with training data.
      5. Developing better test cases for Sayegh.
      6. Updating abcDE to support manual segmentation.
      7. Completing and polishing abcD for entire Beringer corpus.
      8. Defining initial benchmark corpora and evaluation methodology.
      9. Implementing convenience methods for reporting benchmark results.
      10. Moving Beringer corpus to MySQL database.
      11. Enhancing Parncutt, following published techniques and pushing beyond them.
      12. Enhancing Hart and Sayegh to return top n solutions.
      13. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
      14. Adding support to abcDE for annotating phrase segmentation.
      15. Debugging Dactylize 88-key circuit.
      16. Collecting fingering data from JB performances in Elizabethtown.
      17. Completing Dactylize II circuit.
      18. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
      19. Defining procedure for sanity test of production automatic data collector (including Beringer data).
      20. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
      21. Implementing end-to-end machine learning experiment, using Beringer abcD data.
      22. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

      Wednesday, June 13, 2018

      2018-06-13 status

      Done

      Methodology

      • Contemplated a survey to find ground truth for interchangeable digits. (Discussed briefly with AFL and JB over email.)

      Model Building

      • Implemented trigram nodes in networkx for revamped Parncutt.

        Doing

        1. Learning Cytoscape for graph visualization to help debug graph code.
        2. Fixing cost functions in Parncutt.
        3. Validating Parncutt cost functions for left hand.
        4. Adding mechanism to learn weights for Parncutt rules from training data.
        5. Reading some fingering pedagogy (C. P. E. Bach, Couperin, Rami Bar-Niv) to develop a vocabulary for talking to pianists.

        Struggling


        1. A body at rest tends to stay at rest.
        2. Does this topic make sense for my new career situation?


        In Scope

        1. Implementing crude automatic segmenter.
        2. Developing staccato/legato classifier.
        3. Demonstrating improved Parncutt via #1 and #2.
        4. Debugging Sayegh model, which produces results inconsistent with training data.
        5. Developing better test cases for Sayegh.
        6. Updating abcDE to support manual segmentation.
        7. Completing and polishing abcD for entire Beringer corpus.
        8. Defining initial benchmark corpora and evaluation methodology.
        9. Implementing convenience methods for reporting benchmark results.
        10. Moving Beringer corpus to MySQL database.
        11. Enhancing Parncutt, following published techniques and pushing beyond them.
        12. Enhancing Hart and Sayegh to return top n solutions.
        13. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
        14. Adding support to abcDE for annotating phrase segmentation.
        15. Debugging Dactylize 88-key circuit.
        16. Collecting fingering data from JB performances in Elizabethtown.
        17. Completing Dactylize II circuit.
        18. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
        19. Defining procedure for sanity test of production automatic data collector (including Beringer data).
        20. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
        21. Implementing end-to-end machine learning experiment, using Beringer abcD data.
        22. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

        Friday, March 23, 2018

        2018-03-22 status

        Done

        Administrivia

        • Requested suspension of TAA support until next January (for tax purposes).
        • Answered interview questions for College of Engineering story on Fifty for the Future award.

        Model Building

        • Completed Sayegh implementation.
        • Learned how to modify TatSu AST. 
        • Implemented "pivot alignment" evaluation method.
        • Refactored code for better reuse in new models, especially for segmenting input.
        • Fixed how first and last fingering were being constrained to ensure model preferences for second and penultimate fingerings were not ignored.
        • Drafted ISMIR abstract.
        • Drew a diagram of the subproblems in the domain.

          Doing

          1. Writing up what we have done so far.
          2. Reimplementing Parncutt model in framework using networkx.
          3. Developing better test cases for Sayegh.
          4. Implementing crude automatic segmenter.
          5. Updating abcDE to support manual segmentation.
          6. Completing and polishing abcD for entire Beringer corpus.
          7. Defining initial benchmark corpora and evaluation methodology.
          8. Implementing convenience methods for reporting benchmark results.

          Struggling

          1. Sayegh model produces results that do not seem consistent with training data provided.

          In Scope

          1. Moving Beringer corpus to MySQL database.
          2. Enhancing Parncutt, following published techniques and pushing beyond them.
          3. Enhancing Hart and Sayegh to return top n solutions.
          4. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
          5. Adding support to abcDE for annotating phrase segmentation.
          6. Debugging Dactylize 88-key circuit.
          7. Collecting fingering data from JB performances in Elizabethtown.
          8. Completing Dactylize II circuit.
          9. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
          10. Defining procedure for sanity test of production automatic data collector (including Beringer data).
          11. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
          12. Implementing end-to-end machine learning experiment, using Beringer abcD data.
          13. Submitting papers to TISMIR. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

          Thursday, March 22, 2018

          2018-03-01 status

          Done

          Administrivia

          • Talked through a few challenges with the NLP Lab.
          • Did a little forum shopping. As a backup plan for ISMIR conference, the TISMIR journal is accepting submissions. 

          Model Building

          • Completed "reentry" evaluation method for strike fingers.
          • Tested edge cases for evaluation and advising methods.
          • Rejected TensorFlow for Sayegh implementation. We just need a trellis graph, a nasty for loop for training, and Viterbi.
          • Stubbed in support for phrase segmentation in modeling framework.
          • Implemented Sayegh training algorithm.
          • Implemented methods to store and recall trained models for reuse.

            Doing

            1. Implementing Sayegh trellis-graph model from scratch, using Python's networkx.
            2. Defining initial benchmark corpora and evaluation methodology.
            3. Implementing convenience methods for reporting benchmark results.
            4. Completing and polishing abcD for entire Beringer corpus.

            Struggling

            1. The otherwise slick parser module I am using (TatSu) produces an immutable AST. This is cramping my style and promises to get worse as we move along.
            2. The Parncutt code is a disaster under Python 3. Lot of rework needed here.

            In Scope

            1. Reimplementing Parncutt model in framework using networkx.
            2. Moving Beringer corpus to MySQL database.
            3. Enhancing Parncutt, following published techniques and pushing beyond them.
            4. Enhancing Hart and Sayegh to return top n solutions.
            5. Re-weighting Parncutt rules using machine learning and TensorFlow. (This seems like a good fit.)
            6. Adding support to abcDE for annotating phrase segmentation.
            7. Debugging Dactylize 88-key circuit.
            8. Collecting fingering data from JB performances in Elizabethtown.
            9. Completing Dactylize II circuit.
            10. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
            11. Defining procedure for sanity test of production automatic data collector (including Beringer data).
            12. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
            13. Implementing end-to-end machine learning experiment, using Beringer abcD data.
            14. Submitting papers to ISMIR 2018. Abstracts due March 23. Papers due March 30. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

            Friday, February 16, 2018

            2018-02-18 status

            Done

            Administrivia

            • Found ISMIR LaTeX template.

            Model Building

            • Modified music21 code base to ignore ornaments and had pull request accepted.
            • Moved to the alpha release of music21 (with some trepidation)
            • Set up grammar, parser, and convenient Abstract Syntax Tree (AST) for fingering language (abcDF) in Python.
            • Leveraged new parser to implement evaluation methods.
            • Implemented Hamming evaluation method for strike fingers.
            • Implemented "natural" evaluation method for strike fingers.
            • Implemented "pivot" evaluation method for strike fingers.
            • Implemented first (striking) finger constraint for Hart model.
            • Drafted "reentry" evaluation method for strike fingers.

              Doing

              1. Debugging and testing "reentry" evaluation method. 
              2. Creating more test cases for "infrastructure" code.
              3. Studying TensorFlow paradigms for connectionist models.
              4. Implementing Sayegh  model (via TensorFlow or from scratch).

              Struggling

              1. The otherwise slick parser module I am using (TatSu) produces an immutable AST. This is cramping my style and promises to get worse as we move along.
              2. The Parncutt code is a disaster under Python 3. Lot of rework needed here.
              3. Generally pulling hair out debugging the "reentry" code, which should be trivial. I am doing something stupid.

              In Scope

              1. Re-implementing Parncutt model in framework. (The graph class I was using does not seem to exist in Python3, so this needs to be reworked.)
              2. Debugging Dactylize 88-key circuit.
              3. Collecting fingering data from JB performances in Elizabethtown.
              4. Completing Dactylize II circuit.
              5. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
              6. Creating abcD for complete Beringer corpus.
              7. Moving Beringer corpus to MySQL database.
              8. Enhancing Parncutt, following published techniques and pushing beyond them.
              9. Defining procedure for sanity test of production automatic data collector (including Beringer data).
              10. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
              11. Implementing end-to-end machine learning experiment, using Beringer abcD data.
              12. Submitting papers to ISMIR 2018. Abstracts due March 23. Papers due March 30. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

              Friday, January 26, 2018

              2018-01-26 status

              Done

              Model Building

              • Refactored Python algorithm ("Dactyler") architecture.
              • Implemented unit test framework.
              • Migrated Hart algorithm to new framework.
              • Enhanced Hart model to support repeated notes with same pitch.
              • Put announced ISMIR 2018 deadlines (March 23 and March 30) on NLP calendar.

                Doing

                1. Revamping my corpus module to deal with multiple voices and polyphony in abc/abcD input.
                2. Supporting first-note fingering constraint in Hart model to enable the "auto-correcting" or "re-entrant" evaluation method suggested by CR.
                3. Implementing DValuation base class to support evaluation methods.
                4. Implementing DHamming edit-distance evaluation method.
                5. Implementing DNatural edit-distance method.
                6. Implementing DPivot edit-distance method.
                7. Implementing DReEntry evaluation method.

                In Scope for Semester

                1. Re-implementing Parncutt model in framework. (The graph class I was using does not seem to exist in Python3, so this needs to be reworked.)
                2. Debugging Dactylize 88-key circuit.
                3. Collecting fingering data from JB performances in Elizabethtown.
                4. Implementing Sayegh model.
                5. Completing Dactylize II circuit.
                6. Developing method to align performance data with symbolic data. I think this is going to be essential if we are to use Dactylize data moving forward and a key part of its proof of concept. I plan to have something for this at the ISMIR demo session (September 22 deadline).
                7. Creating abcD for complete Beringer corpus.
                8. Moving Beringer corpus to MySQL database.
                9. Enhancing Parncutt, following published techniques and pushing beyond them.
                10. Defining procedure for sanity test of production automatic data collector (including Beringer data).
                11. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
                12. Implementing end-to-end machine learning experiment, using Beringer abcD data.
                13. Submitting papers to ISMIR 2018. Abstracts due March 23. Papers due March 30. Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models, demo of method to align collected performance data with symbolic score.

                Struggling

                1. music21 support for abc is either buggy, or I can't read. Having a hard time splitting an abc file into right- and left-hand parts, something that should be trivial.
                2. The Parncutt code is a disaster under Python 3. Lot of rework needed here.
                3. Contemplating using Hart as one of the two proof-of-concept models for the methodology paper. I might even be planning on it at this point.

                Saturday, December 30, 2017

                2017-12-31 status

                Done

                Administrivia

                • Secured $3000 Elizabethtown College Faculty Grant award with JB in late April.
                • Successfully defended prelim proposal in June.
                • Got a full-time job in September at UIC CCTS.

                Data Collection

                • Repaired original Dactylize circuit damaged in move from NLP Lab.
                • Installed Dactylize system in JB's studio at Elizabethtown College in October.
                • Prototyped new glove assembly for Dactylize.
                • Obtained parts for repair and new Dactylize II system with grant funds and started circuit assembly.
                • Successfully loaded Survey III data from Qualtrics to new MySQL diii2 schema.
                • Loaded missing Survey I and II data in MySQL to new didactyl2 schema, so we can now perform queries across both data sets, as originally envisioned, and all of our fingering data are now in one place.
                • Defined procedure for data collection from performances using GitHub.

                  Doing

                  1. Updating Didactyl framework to recognize abcD format.
                  2. Completing Parncutt model implementation in framework.
                  3. Debugging Dactylize 88-key circuit.
                  4. Collecting fingering data from JB performances in Elizabethtown.
                  5. Implementing Sayegh model.
                  6. Implementing "autocorrecting" evaluation method suggested by CR to compare models head to head.

                  In Scope

                  1. Completing Dactylize II circuit.
                  2. Creating abcD for complete Beringer corpus.
                  3. Enhancing Parncutt, following published techniques and pushing beyond them.
                  4. Defining procedure for sanity test of production automatic data collector (including Beringer data).
                  5. Defining corpora for Dactylize data collection (WTC, Beringer, ??).
                  6. Implementing end-to-end machine learning experiment, using Beringer abcD data.
                  7. Submitting papers to ISMIR 2018 (deadline not yet announced, probably in March). Ideas: a follow-up demo paper describing Dactylize data collected; a full-length paper describing application of evaluation method to models developed; a full-length description of enhanced and/or novel models.

                  Struggling

                  1. Getting reacquainted with Python and music21.

                  Wednesday, February 8, 2017

                  2017-02-08 status

                  Done

                  Data Collection

                  • Created initial table of contents for prelim proposal.
                  • Debugged LaTeX conflict between color package and uicthesi class. Short answer: don't use color.

                    Doing

                    1. Drafting early chapters of prelim proposal.
                    2. Drafting Elizabethtown College grant proposal, dove-tailing with Dactylize chapter in prelim proposal.
                    3. Strategizing initial constraints for system--start with finger legato? Outlining planned evolution of model.
                    4. Strategizing enhanced similarity measure. Minimum edit distance with all (weighted) substitutions? Pivot detector?
                    5. Legato detector? "Phrase" detector?
                    6. Debugging, and defining test cases for, post-processing code (dactylizer.pl).

                    Back-burnered

                    1. Creating abcD for complete Beringer corpus.
                    2. Defining procedure for sanity test of production automatic data collector (including Beringer data).
                    3. Defining procedure for initial data collection sessions (including Beringer data).
                    4. Implementing end-to-end machine learning experiment, using Beringer abcD data.

                    Struggling

                    • So much grading, so little time.
                    • Refinancing house?
                    • Get a job, hippie.

                    Wednesday, October 5, 2016

                    2016-10-05 status

                    Done

                    Administrivia

                    • Massaged resume.
                    • Interviewed with Bank of America Merrill Lynch. It went okay, I think.

                    Data Collection

                    • Conducted two initial automatic data collection sessions with advanced pianist.
                    • Identified key wiring connection problems, which we are attributing to non-conductive adhesive in aluminum foil tape.
                    • Removed tape connections to 48 keys and cleaned key surfaces with Goo Gone.
                    • Applied new copper tape substrate to posterior section of aluminum tape on white keys.
                    • Rewired 48 keys with soldered connections to copper tape subsequently physically attached (taped) to copper substrate with conductive adhesive.
                    • Identified further problems that appear to be primarily post-processing (software) issues.

                       Doing

                      1. Ordering fans and coffee for the lab from Amazon. Deliver to Department?
                      2. Assembling materials for Chancellor's Graduate Research Award.
                      3. Defining test cases for post-processing code (dactylizer.pl).
                      4. Drafting email to invite Survey II subjects to participate in on-campus sessions.
                      5. Creating abcD for complete Beringer corpus.
                      6. Defining procedure for sanity test of production automatic data collector (including Beringer data).
                      7. Defining procedure for initial data collection sessions (including Beringer data).
                      8. Implementing end-to-end machine learning experiment, using Beringer abcD data.

                      Struggling

                      • Need to work on my interviewing skills, which are quite rusty.

                      Wednesday, September 28, 2016

                      2016-09-28 status

                      Done

                      Administrivia

                      • IRB approved protocol amendment #4.
                      • Submitted $400 ISMIR 2016 registration for reimbursement from BDE's ICR fund.

                      Data Collection

                      • Confirmed automatic data collection Oct 1 at 10 a.m.
                      • Transcribed left hand notation for arpeggios and broken chords in Beringer (technical exercise) corpus.
                      • Produced abcD to capture right and left hand fingerings for arpeggios and broken chords in Beringer corpus.
                      • Produced abcD to capture right and left hand fingerings for major and minor scales in Beringer corpus.
                      • Mounted Dactylize circuit boards on back of piano and removed excess wiring. 
                      • Constructed more finger assemblies.
                      • Performed initial sanity test with beginning and intermediate pianists.
                      • Reproduced "slow start" problem in Dactylize system, where initial 50 or so notes in "monitor" sessions all have large numbers of missing or incorrect fingerings assigned. Then the system stabilizes and behaves perfectly. This needs to be debugged, but for the first session, we will just include a 50-note "preamble" for each session. 

                         Doing

                        1. Assembling materials for Chancellor's Graduate Research Award.
                        2. Drafting email to invite Survey II subjects to participate in on-campus sessions.
                        3. Creating abcD for complete Beringer corpus.
                        4. Defining procedure for sanity test of production automatic data collector (including Beringer data).
                        5. Defining procedure for initial data collection sessions (including Beringer data).
                        6. Implementing end-to-end machine learning experiment, using Beringer abcD data.

                        Struggling

                        • Focus has been difficult this week, as I was laid off from my job on Monday after almost 17 years. The way forward is suddenly somewhat dim.

                        Tuesday, September 13, 2016

                        2016-09-12 status

                        Done

                        Administrivia

                        • Submitted protocol amendment #4 to IRB on Sep 7. No word yet.

                        Data Collection

                        • First automatic data collection scheduled for Oct 1.
                        • Transcribed left hand notation for Beringer (technical exercise) corpus.
                        • Transcribed harmonic minor scales for Beringer corpus.
                        • Implemented Dactylize circuit on breadboards for debugging.
                        • Constructed more finger assemblies.

                           Doing

                          1. Assembling materials for Chancellor's Graduate Research Award.
                          2. Drafting email to invite Survey II subjects to participate in on-campus sessions.
                          3. Contemplating supporting multiple "tunes" in abcD. This would complicate abcDE.
                          4. Creating abcD for Beringer corpus.
                          5. Defining procedure for sanity test of production automatic data collector (including Beringer data).
                          6. Defining procedure for initial data collection sessions (including Beringer data).
                          7. Implementing end-to-end machine learning experiment, using Beringer abcD data.

                          Struggling

                          • Waiting for IRB approval to pull data from Survey III (WTC).

                          Monday, June 13, 2016

                          Dactylize build photos

                          I set up a Google Photo album to track progress on the build. Enjoy as order is made out of chaos.

                          Saturday, June 11, 2016

                          Dactylize parts list

                          Here are the items (and approximate costs) needed to reproduce the system I built, assuming you are working in a well-found electronics lab. These parts should be sufficient to produce the piano component and 6 hand assemblies (enough for three people):
                          • A digital piano. (I used an aging Casio Privia PX-130, which I would value at about $300. The latest of this line of instruments, the PX-160, goes for about $500 new.)
                          • Arduino Micro Pro equivalent. (Clones like those from Osoyoo can be had for about $9.) Any Arduino should do the job.
                          • Two 74HC165 8-bit input shift register integrated circuit chips. Cost: $5.
                          • Eleven 74HC595 8-bit output shift register integrated circuit chips like these. Cost: $10.
                          • Eighteen carbon film 10k Ohm 1/4 watt resistors. Cost: $2.
                          • 2 monolithic 0.1 uF 100 Volt capacitors. Cost: $1.
                          • Through-hole protoboard for input shift register (finger) circuit. I used one of these. Cost: $5.
                          • Large solderable breadboard like this for the output shift register (key) circuit. Cost: $12.
                          • Four 40-position male header pins to be cut to various lengths (including eleven 8-pins for key inputs and four 4-pins for finger inputs, plus miscellaneous others). Cost $2.
                          • Roll of 2-inch aluminum duct-work tape. Cost: $5.
                          • Roll of 1-inch copper foil tape for black keys (optional). Cost: $15.
                          • 88 12-inch male-female stranded jumper wires with 0.1 inch pre-crimped terminals. Sold in bags of 50. Cost: $36.
                          • 32 36-inch female-female stranded jumper wires with 0.1 inch pre-crimped terminals. Sold in bags of 20. Cost: $36.
                          • 86 (56 for keys + 30 for hand assemblies) 6-inch male-female stranded jumper wires with 0.1 inch pre-crimped terminals. Sold in bags of 50. Cost: $20.
                          • 45 (11 + 4+4 + 7+7 + 12) 1x8-pin crimp connector housings. Add 2 such housings for each hand assembly you wish to, well, assemble. Get 5 bags of 10 to be safe. Cost: $5.
                          • 12 1x4-pin crimp connector housings. Cost: $2.
                          • 12 1x1-pin crimp connector housings. Cost: $1.
                          • 25 feet of TechFlex Flexo braided cable sleeve (optional). Cost: $10.
                          • Two 10-foot 16-conductor ribbon cables (or about 16 feet of the stuff). Cost: $10.
                          • 10-foot 10-connector ribbon cable. Cost: $5.
                          • 14 16-position right-angle female cable-mounted IDC connectors. Cost: $6.
                          • Six-inch cable zip ties, pack of 100. Cost: $2.
                          • Adhesive cable tie bases, pack of 100. Cost: $10.
                          • 40 latex finger cots. Cost $3.
                          • Three ounces DAP Wedgwood contact cement. Cost: $9.
                          • Graphite lubricant. Cost: $6.
                          • 6-foot USB-to-micro-USB cable. Cost: $8.
                          • 6-foot USB 2.0 A-male-to-B-male cable (or equivalent MIDI cabling for your digital piano). Cost: $5.
                          The total cost for materials is $271 plus the cost of the piano, which you should be able to restore to its original playing condition when finished with your Dactylize system.

                          If you are starting completely from scratch, as I was, you will need the following essentials.
                          • Multi-colored solid hook-up wire. Cost: $18.
                          • Soldering iron. I strongly recommend the Hakko FX-888D soldering station. My experiences as a novice with cheap irons were extremely unpleasant. Cost: $100.
                          • Solder. $10.
                          • Solder wick. $7.
                          • Solder sucker. $5.
                          • Wire strippers. $11.
                          • Needle-nose pliers.
                          • Small wire cutters.
                          • Auto-ranging digital multi-meter that beeps for connectivity testing. The Innova 3320 seems to be the cheapest option that ticks all of the boxes. $26.
                          • Safety glasses.
                          • Flat wooden board to work on.
                          • Electrical tape.
                          • Ice packs for burns.
                          Also recommended:

                          Friday, May 27, 2016

                          Hooah, Dactylize

                          Output shift register

                          Moving directly to the patch-bay matrix circuit turned out to be overreaching. So I stepped back and tested the components separately. First, we made sure the output shift registers were doing the right thing. After some gnashing of teeth, I realized I was loading the wrong sketch. With the right sketch, it worked like a champ. All 16 LEDs lit up in order.

                          Output shift register circuit with full array of test LEDs.
                          See pp. 436-441 in Drymonitis.

                          Input shift register

                          The real problems were in the input shift register. A resistor had a bad connection. With that fixed, its test circuit worked.

                          Input shift register circuit with push buttons for testing through Pure Data.
                          See pp. 430-436 in Drymonitis.

                          Patch-bay matrix

                          Putting  these bits together, we get a working 16-input, 16-output patch-bay matrix. This sometimes gets a little noisy, with false connections detected and then usually corrected. But restarting the patch seemed to quiet it down.

                          Working Broken patch-bay matrix circuit.

                          The plan now, for the first implementation at least, is to have two input shift register ICs (for the hands) and 11 output shift register ICs (for the keys). The Fritzing diagram shared earlier has this the other way around. The output circuit is simpler and requires no resistors or capacitors. This will shorten the part list. I am not sure which way is more performant. I think removing the capacitors will help performance, but it probably won't matter much.

                          Finger-key interface (using mini keyboard)

                          Ignoring scalability, the last bit in the hardware proof of concept is to prototype (however crudely) the patch cord to be formed by connecting a wire emanating from a conductive surface on a given key to one attached to an individual finger.

                          Solder 16x16 patch-bay matrix circuit

                          [WE ARE (still) HERE ]

                          I am a bit concerned about noise in the circuit. Solid connections should mitigate some of the noise, I hope. Also, the dimensions of the separate power supply will not fit on the narrower breadboard-like solderable circuit boards I have purchased. We could leave it on a breadboard, I suppose. I was also surprised to see its LED light up dimly when it was connected to the patch-bay matrix. I took it out in the picture above because I was afraid this phenomenon was contributing to the noise in the circuit. We will need to reintroduce this in the soldered circuit, as I don't think the Arduino can power the complete circuit. But I think this power supply (with the 2 amp 7 volt DC adapter I am using) can power the whole shebang.

                          UPDATE 2016-06-01
                          We have soldered two boards to implement a 16x48 patch bay:


                          The power supply won't fit on the protoboards, so we keep it on a breadboard:


                          The more complicated input shift registers are used for the fingers:


                          Output shift registers handle the key connections:


                          And the last piece is the Arduino UNO:

                          Unfortunately, the "noise" we noted earlier and hoped would be mitigated with a more solid circuit seems instead to be amplified. Or something isn't soldered correctly. If the state of the connections does not change, the output should not change. But we are seeing a lot of output.

                          I have moved back to the breadboard and will step through the software "sketch" to see what I can see.

                          UPDATE 7 June 2016
                          We have a working patch-bay matrix for one input chip and three output. There is apparently something wrong with the daisy chaining of the input shift register component. The plan is now to implement something end to end for one hand.

                          Debug the sketch

                          UPDATE 7 June 2016
                          We have a working sketch for one hand that produces human readable serial output that seems to reflect fingerings accurately.

                          MIDI processing and event time-stamping

                          The idea is to monitor the following:
                          • Fingering connection change events
                          • MIDI note-on and -off events
                          • MIDI program change events
                          We will do this through Pure Data. Each event shall be timestamped at microsecond granularity. (We will then process this output offline with a Perl script to align the fingering events with the MIDI events and produce some flavor of abcDF. This programming bit will probably be done after most of the hardware and other software is assembled and working. Initially, we just need to convince ourselves that we have all of the data we need to get this to work.)

                          UPDATE 7 June 2016
                          I punted on Pure Data. I have a multi-threaded Python script that monitors the Arduino serial output and the MIDI input and applies a microsecond timestamp uniformly to both. This should be enough for us to post-process the data to determine fingerings. The exact logic of how that should be done is left as an exercise for the writer.

                          Foil's war

                          Now we move to the full size keyboard. The plan for applying foil to the keys was to use copper foil tape for the white keys and aluminum foil tape for the black keys.

                          To make the connection from the foil to a (stranded) copper wire, I planned to solder the wire to a small piece of tape and then stick this tape on the key foil. This seems like a potential problem area as we are relying on the adhesive to make this connection with a moving part. Soldering directly a wire directly on the foil either before or after might be a better approach.

                          It might also be more effective to use thicker metal foil and apply contact cement ourselves. This might help us position precut (and pre-soldered?) foil pieces, as the cement will take some time to dry.

                          Another idea is to crack open the keyboard enclosure and move the connection (and maybe our entire circuit) inside the enclosure. This would be elegant, but we don't know how this will fit. And, while this does seem to protect us from accidental bumps and things getting unstuck, fixing any problem that does occur will be a time-consuming affair. Just a thought.

                          Do we need to use vegetable oil to solder the wire to the foil?

                          We start foiling 16 keys for the initial 16x16 circuit.

                          UPDATE 7 June 2016
                          We have foiled about 3/4 of the keys. Here are some pictures




                          A feast of wire (crimping my style)

                          We create ribbon cables and connectors that allow us us to reduce strain on the key wire connections and foil adhesive and easily to detach the keys and the finger assemblies from the circuit. Start with end to end solution for the 16x16 (10x16) circuit.

                          All connections to key foil and finger cots should be with stranded wire. The other ends of finger wires will feed 4-wire crimp connectors. In the likely event that our circuit will be housed outside the keyboard's own enclosure, the other ends for key wires will feed 8-wire crimp connectors, which will be firmly attached to the outside of the keyboard enclosure to reduce potential strain on the foil connections.

                          The original plan was to use banana plugs for individual finger connections to the circuit board. This will be abandoned in favor of crimp connectors interfacing with IDC connections and ribbon cables.

                          UPDATE 7 June 2016
                          We are booting on the ribbon cables for the key connections. We plan to mount the boards on the back of the music stand and connect directly to it using Pololu jumper connector housings. We don't need long run ribbons. Here is a picture of what we plan to do for the key harnesses:


                          Proper finger harness

                          Latex cots need to have conductive material (either foil or graphite-infused contact cement) applied and attached to (flexible multi-stranded) wires. The wires need to be collected at the wrist with medical tape . Also at the wrist, crimp connectors should allow for easy attachment to (and detachment from) ribbon cables. These ribbon cables will be connected to the male header pins on the  input shift register component.

                          This is probably going to take some iteration/experimentation to get right.

                          UPDATE 7 June 2016
                          Here are two options for the end to-end finger connections, one which requires soldering and one which does not.

                          Soldered finger connection assembly.

                          Solderless finger connection assembly.

                          Add one output shift register IC at a time and test

                          Yeah, what he said.