WEBVTT

1
00:00:03.750 --> 00:00:19.770
Christopher Jordan: Okay, thanks everyone for joining us. This is our second session of kind of introducing and continuing the ISO bank ingest and now interface testing process.

2
00:00:20.670 --> 00:00:34.770
Christopher Jordan: Several of you have already provided us with some sample data and everything. Everything thus far that we have tried to ingest the provided has we've been able to do so.

3
00:00:35.280 --> 00:00:48.660
Christopher Jordan: Sometimes with some tweaks and needed and sometimes learning something in the process. So, so thank you especially to everybody who has sent us stuff already. It's been extremely helpful.

4
00:00:49.830 --> 00:01:03.570
Christopher Jordan: I'm going to introduce today by talking just a little bit more about some of what we have some of what we've learned in some of what has changed as a result of the initial

5
00:01:04.980 --> 00:01:17.340
Christopher Jordan: feedback that we've gotten from you all in the process of learning how to ingest data that that comes from the real world. And then I'm going to pass over to David walling

6
00:01:18.090 --> 00:01:31.170
Christopher Jordan: Who's going to kind of walk you through the template generation and the ingest process as it is today. And that's going to act as kind of an introduction to to the process that we hope you will all

7
00:01:31.650 --> 00:01:42.240
Christopher Jordan: I'll be able to use and test yourselves in the upcoming weeks and before I go into updates. I want to just take a moment to acknowledge the

8
00:01:42.690 --> 00:01:52.290
Christopher Jordan: The fantastic work that that David and Thomas love and and have all done on putting together the the technical foundation

9
00:01:52.920 --> 00:02:03.360
Christopher Jordan: For the metadata model and for actually inputting all of the metadata into the system, as many of you know we have a an extraordinarily

10
00:02:03.810 --> 00:02:15.270
Christopher Jordan: Rich and robust metadata model with literally hundreds of of elements and we've come up with a system that will allow us to support a fairly flexible.

11
00:02:15.870 --> 00:02:25.230
Christopher Jordan: Model with changing that data fields and changing controlled vocabularies over time without having to do significant reengineering but

12
00:02:25.830 --> 00:02:35.250
Christopher Jordan: One of the consequences of that is that Anna has had to manually create all these metadata fields and be controlled vocabularies

13
00:02:35.910 --> 00:02:52.860
Christopher Jordan: In the, in the interface and and she's, she's done a ton of work on that as Tomislav and David have done on the technical foundation so I'm really, I'm really grateful to have had them to do all of that because I, I certainly wasn't smart enough to come up with it so

14
00:02:54.840 --> 00:03:02.070
Christopher Jordan: Okay, so a couple of things that we learned in the in the process of kind of ingesting

15
00:03:03.240 --> 00:03:14.580
Christopher Jordan: The data that we've gotten from you already one of the first and kind of most important things we had initially been using a combination of the

16
00:03:15.240 --> 00:03:27.150
Christopher Jordan: The user identity and the sample ID that the user provided as a way to uniquely identify both samples and measurements. We also had a lab sample ID.

17
00:03:27.990 --> 00:03:39.690
Christopher Jordan: That was meant to be, you know, for a laboratory provided identity on the sample, but it kind of quickly became clear that we also had these cases where we have

18
00:03:40.620 --> 00:03:52.380
Christopher Jordan: Multiple different analyses happening on the same sample. And so we need to be able to uniquely identify both a sample provided by a user and the

19
00:03:53.280 --> 00:04:11.790
Christopher Jordan: multiple measurements that are associated with that sample. So as a result of that we we elevated. We already had an analysis identifier in the in the QA QC metadata, but we elevated that up to the core metadata and it's now a required field.

20
00:04:12.840 --> 00:04:25.860
Christopher Jordan: So, you know, many, many folks I think will not have a an analysis ID that's provided by the lab, either because the lab didn't provide it or because the day that you're submitting

21
00:04:26.670 --> 00:04:38.760
Christopher Jordan: Just you. You don't have access to that level of of meta data. And that's, that's fine. It doesn't have to be a really precise value. It just has to be something that is unique to that sample.

22
00:04:40.320 --> 00:04:53.340
Christopher Jordan: So that we can, if necessary, differentiate different kinds of analyses that produce different kinds of measurements, but coming from the same sample that you provided to the lab.

23
00:04:54.270 --> 00:05:07.500
Christopher Jordan: So that's a it's it's shouldn't be a very consequential change in terms of filling out and providing the data to us, but it's, it's something that you need to be aware of.

24
00:05:08.100 --> 00:05:17.250
Christopher Jordan: Since you know it may, it may differ from your expectations of what the required fields are based on kind of previous discussions workshops

25
00:05:18.420 --> 00:05:31.110
Christopher Jordan: The other kind of just minor thing we we have a number of fields in the place where this came up most frequently was the reference materials.

26
00:05:31.740 --> 00:05:45.630
Christopher Jordan: For the lab and analysis metadata, where the the field can have multiple values in it. And rather than asking you to do something within a comma separated

27
00:05:46.800 --> 00:05:57.480
Christopher Jordan: Spreadsheet to somehow differentiate multiple values within a comma separated spreadsheet. What we're asking folks to do is is just duplicate the column.

28
00:05:58.440 --> 00:06:04.440
Christopher Jordan: And in that case, and provide us multiple columns with the primary reference material or whatever.

29
00:06:04.980 --> 00:06:20.040
Christopher Jordan: Whatever field that may be, and we can handle that appropriately as long as we know that the metadata is something that we can have multiple values for that's something that the system supports. We have talked a little bit about

30
00:06:21.330 --> 00:06:32.160
Christopher Jordan: Maybe providing more differentiation there. So, allowing for a primary reference material one primary reference material to and so on so forth. Because some systems.

31
00:06:32.790 --> 00:06:48.870
Christopher Jordan: We've, we've heard that some of the data wrangling systems don't like having multiple columns with the same header on them. So that's kind of silly and I know been topic of discussion. But for now, that's, that's the guidance and that's that's the way that the system.

32
00:06:50.880 --> 00:06:57.780
Christopher Jordan: Works with the data and that will hold true for for fields, the GPU generate through the, through the template generator

33
00:07:00.990 --> 00:07:11.070
Christopher Jordan: And then finally, one something else I wanted to just highlight from the work we've done so far is we do have in the system for each of the

34
00:07:11.640 --> 00:07:24.780
Christopher Jordan: metadata fields. We also have descriptive information we have descriptive information both for the metadata field itself. And in the case where there's a controlled vocabulary for each of the

35
00:07:25.740 --> 00:07:33.390
Christopher Jordan: vocabulary terms within the list that we provide. And as long as we have those descriptions, we have a nice and

36
00:07:34.590 --> 00:07:41.400
Christopher Jordan: If it's available. I think David will show you this. We have a nice kind of Field Guide to the metadata that will explain to you.

37
00:07:42.120 --> 00:07:42.870
Christopher Jordan: What the

38
00:07:43.050 --> 00:07:52.620
Christopher Jordan: What the valid values are what the metadata field is meant to be. And then details about all of the individual controlled vocabulary elements when we have them.

39
00:07:53.310 --> 00:08:00.300
Christopher Jordan: The only consequence of that is right now we have, at best, very minimal descriptions for all of those

40
00:08:00.990 --> 00:08:07.410
Christopher Jordan: That may come from you know from previous discussions in the workshops. But in a lot of cases, we don't have the

41
00:08:08.040 --> 00:08:14.760
Christopher Jordan: Necessary expertise to provide appropriate descriptions so kind of just a heads up for everybody that we will

42
00:08:15.120 --> 00:08:22.530
Christopher Jordan: We will be asking for the Community's help and providing useful descriptions for all of these various metadata fields.

43
00:08:22.920 --> 00:08:36.060
Christopher Jordan: And I'll just say now if you see something as you're working with the system and you you have or no quickly how to write a couple of sentences that describe it do so and send it to us. We can put it in pretty quickly.

44
00:08:37.380 --> 00:08:39.210
Christopher Jordan: If not, then we will be

45
00:08:40.350 --> 00:08:55.470
Christopher Jordan: Organizing kind of a more systematic mechanism to get all of those all of those metadata descriptions populated because it is a very nice feature of the system. And I think it will be. I think it will be very helpful over the long term.

46
00:08:56.940 --> 00:09:10.350
Christopher Jordan: In helping us data depositors to understand what exactly is meant by each of these metadata descriptions. What, what kind of data goes in each one of these fields.

47
00:09:10.980 --> 00:09:18.030
Christopher Jordan: Because this is, as I mentioned before, the, the metadata model is is very complex. There's lots of terms and it's not

48
00:09:18.330 --> 00:09:28.620
Christopher Jordan: It's not possible just from the descriptions of all of the fields to have perfect clarity as to exactly what it what we mean by every single metadata field.

49
00:09:29.010 --> 00:09:37.530
Christopher Jordan: And how that matches up with your data that's kind of another another lesson. I'd say we've taken from the early deposit ingest process.

50
00:09:40.290 --> 00:09:50.790
Christopher Jordan: Okay, so I'll just pause there. If there are, if there are any questions about any of that. I'm happy to take them. Otherwise, we can move on to to David in the demo.

51
00:09:51.750 --> 00:09:55.050
Mariel Lee Campbell: I did have a quick question that I mentioned in my email earlier.

52
00:09:56.100 --> 00:09:58.080
Mariel Lee Campbell: Today I'm sorry I came in late. This is Mario

53
00:09:59.460 --> 00:10:00.420
Mariel Lee Campbell: Can up

54
00:10:02.100 --> 00:10:11.400
Mariel Lee Campbell: Did you mention in earlier discussions, for example, where we would want to put scientific names or is that part of the description that you're just talking about

55
00:10:12.180 --> 00:10:12.480
Text.

56
00:10:13.500 --> 00:10:23.730
Christopher Jordan: Yeah, so, so the the version of the template that you were looking at was the earlier version that that definitely did not have tax on names.

57
00:10:24.420 --> 00:10:34.890
Christopher Jordan: There is there is a space for that in the fuller metadata model and and i think i think you will probably see that, if not, if not today then then very soon.

58
00:10:35.820 --> 00:10:48.000
Christopher Jordan: One thing I will mention about that is that for now. That's going to be a free text field we do intend to do validation against that, in the long term.

59
00:10:48.870 --> 00:10:58.440
Christopher Jordan: But so far, we have good suggestions, but I wouldn't say we have kind of solid consensus as to how exactly we want to do validation of the

60
00:10:58.800 --> 00:11:11.160
Christopher Jordan: The taxonomic names in that field. So we're going to kind of, we're going to kind of punt on it for the time being, and trust that you all mostly know what you're doing as far as inputting taxonomic names and it won't be too far off.

61
00:11:11.910 --> 00:11:14.370
Mariel Lee Campbell: And then one other question in terms of locality.

62
00:11:16.710 --> 00:11:29.910
Mariel Lee Campbell: We spoke about having some mitigate data guidelines for how to input data for these various fields and would be interesting to know like what level of locality. Are you assuming that we would want to include higher geography.

63
00:11:31.020 --> 00:11:38.910
Mariel Lee Campbell: So total calories on the paper that I was putting data in for just said something like mainland, but I know that this is United States Alaska.

64
00:11:39.420 --> 00:11:48.420
Mariel Lee Campbell: In a particular quad. And so I just wondered if there was plans were plans to try to do standardized locality information or if that would also be free text.

65
00:11:50.280 --> 00:11:55.920
Christopher Jordan: He. I mean, for the time being, we're asking more for for coordinates than localities.

66
00:11:56.730 --> 00:12:12.060
Christopher Jordan: The only thing without getting into the longer discussion. The only thing I'll say about about localities is that I would like to support it, but I think you have a similar issue there about word is your where does your validation come from what is, what is your source.

67
00:12:13.320 --> 00:12:18.960
Christopher Jordan: You know, and, and we can we can rely on sources like Archos for a lot of this stuff, but

68
00:12:20.190 --> 00:12:26.400
Christopher Jordan: You know, especially if we have multiple potential locality sources, then, then things get complicated so

69
00:12:27.240 --> 00:12:39.810
Christopher Jordan: It's, it's definitely something I'd like to support, but it's it's a it's a more complicated topic. So for the, for the time being, I think we're more focused on just getting, you know, GPS coordinates, when we, when we can

70
00:12:40.350 --> 00:12:40.980
Mariel Lee Campbell: Okay, great.

71
00:12:45.600 --> 00:12:45.900
Christopher Jordan: Okay.

72
00:12:47.160 --> 00:12:54.630
Christopher Jordan: Alright. So, David, I think you should have the ability to share your screen. Let me know if you need me to change anything.

73
00:12:55.950 --> 00:12:57.030
Christopher Jordan: Okay. Thanks, Chris.

74
00:12:59.070 --> 00:12:59.760
Can you guys hear me.

75
00:13:05.340 --> 00:13:06.930
Christopher Jordan: Yes, we can hear you. Okay.

76
00:13:08.640 --> 00:13:20.070
David Walling: Yeah. So here, we're still working and pushing all of our latest changes to a quality assurance or test website. And this is kind of where we want people to initially actually come in and use the site.

77
00:13:21.390 --> 00:13:26.550
David Walling: This is something will keep around throughout the life of the project. So as we push out new features we can first have a

78
00:13:27.090 --> 00:13:39.300
David Walling: Kind of a friendly user group test those out before we push them all the way up to production. So the URL is just ISO bank dash QA for quality assurance dot t ACC in texas.edu

79
00:13:39.990 --> 00:13:47.490
David Walling: And we're very close to just having you guys come in and register for accounts and just to turn you loose on this website.

80
00:13:48.000 --> 00:13:58.620
David Walling: There's a couple of things we want to do behind the scenes to finish up. That is, we want to make sure our backups are all set up to automatically take snapshots, even in the QA version so that if you do provide test data.

81
00:13:59.190 --> 00:14:09.840
David Walling: We don't lose it. So we need a couple more days to get all that set up, but I did want to first go ahead and walk through very briefly how to register for an account and I so bank.

82
00:14:10.890 --> 00:14:19.920
David Walling: So most of these pages will be open to the general public. So these first three tabs, for instance, our CMS generated

83
00:14:22.080 --> 00:14:23.910
David Walling: public facing pages.

84
00:14:25.050 --> 00:14:31.200
David Walling: Are data sets statistics page which we recently added will can also be something that's just available to the general public.

85
00:14:31.860 --> 00:14:39.660
David Walling: But there's going to be certain features like my data that are actually going to require you to first log in, which of course means, you first need to have an account.

86
00:14:40.140 --> 00:14:49.380
David Walling: So if you click Register or create account, you'll get a simple form, like you see on pretty much any website and I can just fill this out real quick.

87
00:14:57.870 --> 00:15:04.770
David Walling: Once you hit register it should send you through an email your just your standard email confirmation step.

88
00:15:06.030 --> 00:15:08.700
David Walling: You've opened that email accounts.

89
00:15:10.860 --> 00:15:15.420
David Walling: Once you get that email, you should have one additional link to activate your account.

90
00:15:17.550 --> 00:15:22.710
David Walling: Account activated and you can now login. So I'm actually going to use

91
00:15:24.900 --> 00:15:25.830
Some news.

92
00:15:32.760 --> 00:15:38.970
David Walling: So now if I try to access see my data, it'll force me to login for now I'm going to use this test account.

93
00:15:40.950 --> 00:15:53.250
David Walling: And all I will say this in QA as you register for your account and start submitting data. The idea is least for this initial phase is we're going to try to move all of that data and user account information up to production as well.

94
00:15:53.580 --> 00:16:04.920
David Walling: Eventually will be at a point where anything tested in QA will not be expected to stay around only the features and code will get pushed to cue to the production site. Once we get to that point.

95
00:16:07.020 --> 00:16:08.190
When login here.

96
00:16:09.930 --> 00:16:17.970
David Walling: So here is the main view for working with your data and the QA site for ISO bank. Here are three of the

97
00:16:18.990 --> 00:16:21.810
David Walling: Sample Data Sets that we've received back so far.

98
00:16:23.400 --> 00:16:32.730
David Walling: And here's where you basically can just manage all of these submitted data sets, which is the term that we use in the actual code, so refer to basically a CSV file.

99
00:16:34.080 --> 00:16:39.810
David Walling: If I want to actually list all of the records associated with that CSV file. We have a separate page for doing that.

100
00:16:40.950 --> 00:16:57.720
David Walling: This is highly editable. We can figure out exactly how much of the information to show on this page, but it's just a quick guide to let you see you know what records are actually associated with the CSV, I can actually sort on all of this, this information.

101
00:16:59.430 --> 00:17:11.760
David Walling: And then for an individual record and by record, we really mean a line in a CSV file, you can click on the details page for it. And this is kind of the more extensive

102
00:17:12.630 --> 00:17:23.430
David Walling: Information about that particular line in a submitted CSV file. So in this case, this particular line only had a single measurement type of detail. Ah.

103
00:17:24.090 --> 00:17:29.730
David Walling: But this particular value and an unknown for the unit, but it did provide the scale information.

104
00:17:30.300 --> 00:17:40.710
David Walling: These are the measurement specific metadata fields. The rest of the additional metadata is all of the metadata that is shared by all the measurement types for a given line.

105
00:17:41.580 --> 00:17:51.900
David Walling: In a CSV file. And for each of those we are grouping them into certain categories, just to make the management of them easier, since we do have so much information and for any given

106
00:17:53.460 --> 00:17:54.840
David Walling: Record. There may be

107
00:17:56.220 --> 00:18:05.940
David Walling: You know, maybe a single value, maybe 50 values and this page will be dynamically created based on the actual metadata values that are present for this particular record.

108
00:18:11.040 --> 00:18:23.190
David Walling: You can also at any time download a CSV, the actual CSV first version of the file that was submitted and this gets back to how we expect updates to these files to work.

109
00:18:23.760 --> 00:18:29.970
David Walling: Well, as I mentioned that the last time we want Excel or whatever your favorite CSV editor.

110
00:18:30.420 --> 00:18:41.910
David Walling: Is to be that primary editing interface because it already does this very well it's scripted, you could be. I think we already have one user is actually using our to spit out and manage kind of all these

111
00:18:42.570 --> 00:18:53.040
David Walling: CSV files. So we want y'all to keep using the tools that you're already comfortable using. We don't want to try to refill recreate all this type of basic functionality within the application itself.

112
00:18:53.700 --> 00:19:04.380
David Walling: So you could come in here update this this information fix them values add to it, whatever. And then you would simply resubmit that now, right now we don't have an update button.

113
00:19:06.030 --> 00:19:08.520
David Walling: So for now, we could simply delete the file.

114
00:19:10.380 --> 00:19:11.730
David Walling: And then just resubmit it

115
00:19:15.030 --> 00:19:28.440
David Walling: We are requiring there to be a unique title for each one of these files you submit, as well as providing some sort of description just to help you and others who come and looking for data kind of understand what this is.

116
00:19:31.470 --> 00:19:33.300
Choose our file.

117
00:19:34.560 --> 00:19:36.900
Which one actually deleted me double check.

118
00:19:38.430 --> 00:19:39.540
David Walling: Go back to my data.

119
00:19:45.120 --> 00:19:45.660
David Walling: Recreate

120
00:19:53.220 --> 00:20:05.220
David Walling: And once I submit, you'll see the button is grayed out. So I can't hit it again. It's good, it's depending on the size of the file and the number of metadata fields. It may take anywhere from five to 30 seconds to actually load all that data.

121
00:20:07.860 --> 00:20:10.500
David Walling: For each line it's you know it's verifying all these

122
00:20:11.640 --> 00:20:19.770
David Walling: Information about the metadata fields, whether it's required, whether it's required only if another field is present on that particular line, so on and so forth.

123
00:20:20.340 --> 00:20:32.160
David Walling: But at the end of that if it successfully was able to create every single line and that CSV file. It'll take you to the summary page which tells you how many records were actually created

124
00:20:32.520 --> 00:20:43.560
David Walling: And then for each of those that actually give you a list of warnings. So these were things that did not cause the file to be rejected that it's the same thing that we want you to be aware of. So for instance,

125
00:20:43.920 --> 00:20:56.160
David Walling: Each row and this particular CSV has a column called external sample ID, which is not matching up to a known column name for a supported metadata field.

126
00:20:56.790 --> 00:21:12.510
David Walling: So this is a feature we wanted to have this means that you can actually include additional columns in your CSV, that we don't yet support and at some point in the future we may add a metadata field that matches that column name, essentially, and you could simply, you know,

127
00:21:13.590 --> 00:21:20.340
David Walling: resubmit that or eventually we'll have an update field. So you can hit Update pass it along the

128
00:21:21.180 --> 00:21:26.370
David Walling: That data CSV file, instead of just deleting all the records will actually update them on the fly.

129
00:21:26.850 --> 00:21:37.200
David Walling: This will become more important as we actually build additional functionality on top of these submitted data sets. So for instance, maybe as we start sharing these between users and projects.

130
00:21:37.470 --> 00:21:46.950
David Walling: You may want to keep a list of your favorite sets. So in that those cases, we're not going to want to just completely delete the data set and recreate it because we're going to have foreign key References

131
00:21:47.820 --> 00:21:58.350
David Walling: Within this kind of favorites, a concept that we want to keep in the database. So the actual technical details behind that is actually not that difficult. It's just we have something we haven't focused on

132
00:21:59.460 --> 00:22:00.360
Fight at this point.

133
00:22:02.970 --> 00:22:07.470
David Walling: We can see that so so in QA right now, these are the only three data sets. We currently have loaded

134
00:22:08.070 --> 00:22:15.810
David Walling: But if we look at our data set statistics page, we see that even with these three we have 267 unique analysis records.

135
00:22:16.080 --> 00:22:26.220
David Walling: Almost 4000 unique analysis metadata records associated with those 658 measurement records and for those 658 measurement records 790

136
00:22:26.760 --> 00:22:41.760
David Walling: Measurement metadata records, not every measurement type has measurement metadata associated with it. So things like percent carbon or just a value. There are no real metadata records, at least at this time directly associated with that with that measurement tight.

137
00:22:43.230 --> 00:22:49.530
David Walling: But we we expect to fully fleshed this page out with some other interesting things and we are open to suggestions.

138
00:22:49.890 --> 00:22:54.600
David Walling: About what should be here. You can imagine the table that shows counts by measurement tight.

139
00:22:54.900 --> 00:23:08.460
David Walling: Counts by analysis type material type, so on and so forth. So this is, again, kind of a page that we can use to market ISO bank eventually to the wider community and it can be accessible to people who haven't even logged into the system.

140
00:23:12.150 --> 00:23:19.860
David Walling: Any questions about this kind of, oh, so I didn't want to show. So what happens if you try to submit a data set that has an error.

141
00:23:22.470 --> 00:23:25.500
David Walling: So if I select one, let's say,

142
00:23:26.700 --> 00:23:35.250
David Walling: This one, which I have some known errors in and I tried to submit that the feedback mechanism is it will if there's any line in that CSV that has an error.

143
00:23:35.610 --> 00:23:44.400
David Walling: It's going to reject the entire file, but also give you a row by row information about what was wrong with that CSV. So in this one.

144
00:23:45.120 --> 00:23:56.040
David Walling: We saw that no one lab sample idea was not provided with that's required field and also enrolled one we had this d 13 see measurement scale field that had a

145
00:23:56.730 --> 00:24:09.120
David Walling: Value of bit bad so I did not match a lookup value. So measurement scale is a lookup value metadata filter. So the value in that if there is a value.

146
00:24:09.930 --> 00:24:13.710
David Walling: It would try to look that up in the database if it can't find it, it will reject it.

147
00:24:14.070 --> 00:24:22.020
David Walling: So on and so forth. So the feedback mechanism is you'll you'll bring your data in, there's going to be based on one of these upload templates which will dive into next

148
00:24:22.350 --> 00:24:37.110
David Walling: And you'll fill that out will come here try to submit it. If there's errors. We want this to provide enough information to the user that they'll be able to go back and fix those errors in the CSV file and resubmit that

149
00:24:40.980 --> 00:24:47.850
David Walling: Any questions on that before I jump into actually, how do we create an upload template to start filling that out.

150
00:24:52.980 --> 00:24:53.280
Okay.

151
00:24:54.360 --> 00:25:02.790
David Walling: So we showed this off at our last meeting, but we've done a little bit of work to kind of hardening it and a few minor enhancements as well.

152
00:25:04.440 --> 00:25:13.470
David Walling: The first thing we'll see here is that we have added this reset form button. So as you're going through and clicking these. There are so many of them that

153
00:25:14.070 --> 00:25:20.880
David Walling: You may get to a point where you're like, Okay, well I just kind of need to start over. So by hitting reset form, it's basically going to

154
00:25:21.690 --> 00:25:45.210
David Walling: Pre select all the required and recommended metadata fields as they're currently specified in the system. This is also the same mechanism that we expect to use to provide community specific metadata templates. So imagine another button that says water or paleo, we could

155
00:25:46.380 --> 00:25:52.860
David Walling: Have that community decide what are the most important or what are the metadata fields that they really want to be in all

156
00:25:53.700 --> 00:26:08.820
David Walling: data sets that are uploaded and associated with that particular community. And so we can have those fields to click on a button, have all of those particular fields pre selected and then at the bottom, allow them to download that template, fill it out and submit it to us.

157
00:26:10.380 --> 00:26:19.500
David Walling: I think the I think since last time. We've also added this measurement types area. So you have to select at least one of these measurement types in order to download the template.

158
00:26:19.920 --> 00:26:29.100
David Walling: If you try to come down here without doing that, you'll see that the selected measurement types is empty and this button is basically turned off with a message to please select the measurement type

159
00:26:29.970 --> 00:26:42.780
David Walling: So if I come back up. Select the comments, carbon, nitrogen fields. They now show up in the selected measurement types and this button has activated.

160
00:26:45.750 --> 00:26:46.890
David Walling: We do have

161
00:26:50.100 --> 00:26:58.770
David Walling: Within the system relationships between certain fields, trying to remember exactly which ones we've been testing on, I think it was

162
00:26:58.860 --> 00:27:05.670
Tomislav Urban: Click today. I always like to use soil in material type and then soil horizon.

163
00:27:06.840 --> 00:27:10.200
Tomislav Urban: As a dependent field. I hope that

164
00:27:10.680 --> 00:27:11.640
Tomislav Urban: QA. I know it is

165
00:27:12.030 --> 00:27:13.440
Tomislav Urban: On my database, but I saw

166
00:27:14.850 --> 00:27:19.830
Tomislav Urban: A note to try sorry the inorganic version of soil scroll down a little bit more

167
00:27:24.330 --> 00:27:25.410
Tomislav Urban: There we go. That one.

168
00:27:28.740 --> 00:27:29.130
Sorry.

169
00:27:34.380 --> 00:27:34.710
Tomislav Urban: Yeah.

170
00:27:37.170 --> 00:27:39.300
Tomislav Urban: Looks like it's not set up in QA

171
00:27:40.530 --> 00:27:45.600
David Walling: So this is a lot of JavaScript stuff that we're still kind of working through some of the bugs. But things like

172
00:27:52.350 --> 00:27:58.020
David Walling: Blinking now and which ones will trigger the there's a lot more actually in QA than there was even here.

173
00:28:02.970 --> 00:28:13.530
David Walling: Well, anyway, the point is that some of these by by clicking one if there's a relationship that exists, it will automatically select one that's dependent on that. So if I select

174
00:28:13.860 --> 00:28:20.250
David Walling: One of these fields and another one is required if that fill the selected it should automatically implement that logic that can also that can be true.

175
00:28:20.250 --> 00:28:28.230
David Walling: With both the just the metadata field, but also certain valid values within a metadata field could trigger a relationship.

176
00:28:28.260 --> 00:28:29.040
Anna Dabrowski: That says that

177
00:28:29.400 --> 00:28:39.660
David Walling: An additional fields should be provided. If this metadata field has a certain value. So that's a lot of what we've been focused on over the last couple of weeks.

178
00:28:44.700 --> 00:28:49.140
David Walling: But nevertheless. Once you've selected your fields, you should be able to hit the download button.

179
00:28:51.150 --> 00:28:54.630
David Walling: Spit out your CSV, open it in your favorite CSV editor.

180
00:28:56.340 --> 00:29:03.630
David Walling: And here we see the all the fields. The main thing I wanted to point out about this and then I don't think that we've really talked about

181
00:29:04.290 --> 00:29:14.490
David Walling: Previously, is that there's a certain order to how we're putting these spills into this, the CSV files. Basically, it's going to be the core metadata fields.

182
00:29:16.380 --> 00:29:23.850
David Walling: For each group of analysis metadata will provide those columns. And then at the end.

183
00:29:24.930 --> 00:29:30.690
David Walling: For each measurement type, we will provide a column for that measurement tight and if that measurement type has

184
00:29:31.320 --> 00:29:45.420
David Walling: metadata associated with it. Those will always follow directly after that the measurement type. So this is done for two reasons. The first is we want to keep these templates and the CSV files that users will download from ISO make

185
00:29:46.470 --> 00:29:55.920
David Walling: as consistent as possible. So it just kind of helps you. If you're looking for a particular value you kind of start to understand where it would be within the CSV files.

186
00:29:56.310 --> 00:30:04.290
David Walling: This is ordering also provides a way to simplify the ingest logic slightly because there is an implicit relationship between

187
00:30:06.210 --> 00:30:17.610
David Walling: Well, it's technical, but if I have like all the court measurement all the core analysis data, I can actually create that analysis objects and then start tacking on all the analysis metadata to it.

188
00:30:18.030 --> 00:30:24.720
David Walling: And so for those two reasons we are at least at this time making the order of these columns have actual meaning.

189
00:30:26.670 --> 00:30:35.340
Mariel Lee Campbell: I have a question regarding that. So if you wanted to add a new field or column you would need to go back and download a new template.

190
00:30:36.420 --> 00:30:39.600
Mariel Lee Campbell: And rather than just inserting the column with the appropriate header.

191
00:30:39.690 --> 00:30:43.170
David Walling: somewhere and you're fine. Actually, you could do both, actually, it would be okay.

192
00:30:44.520 --> 00:30:55.470
David Walling: So if it was a new analysis level metadata. You could just tack it on you. It's probably wouldn't break the ingest. But again, we do want to try to keep them as

193
00:30:56.700 --> 00:30:57.900
consistent as possible.

194
00:30:59.130 --> 00:31:06.750
David Walling: So you could do it both ways to do this. These don't have to come from the download tool or the generator tool you could

195
00:31:07.200 --> 00:31:22.740
David Walling: Create your spill it out and then hand it to a partner is like an example of doing and have them fill out their own version of that just based on the CSV file we do anticipate that these things could just be passed around as blank templates or you know header or CSV files.

196
00:31:23.430 --> 00:31:33.510
Mariel Lee Campbell: But the order would need to be the same as it would be if it were you downloaded the headers. You couldn't just depend or add in a column in some random place within the file.

197
00:31:35.700 --> 00:31:39.810
David Walling: That's a good point. It may break it. There's, there's, each individual

198
00:31:41.400 --> 00:31:49.800
David Walling: One doesn't have to be the exact same order. It's more the groups of them right. It's more like these first four need to be the need to be in some order at the start.

199
00:31:50.220 --> 00:32:02.100
David Walling: All the analysis metadata needs to come in the middle, followed by the measurement stuff at the end. You could if it's an analysis measurement field, it could go anywhere within this middle group.

200
00:32:03.240 --> 00:32:09.750
David Walling: But that's actually a good point that I'll have to kind of think about the implications of how that might complicated and just

201
00:32:10.410 --> 00:32:18.840
Christopher Jordan: Yeah, I think the the point the point there would just be, you know, if you if you know roughly where it goes and exactly what it's called.

202
00:32:19.200 --> 00:32:34.890
Christopher Jordan: It's probably okay to just add it to the spreadsheet, but but there's not going to be any guarantee if if you are just kind of randomly sticking new fields into the middle of the spreadsheet that didn't work the way you expect it to.

203
00:32:35.700 --> 00:32:37.470
Philip Manlick: The follow up on that, Chris.

204
00:32:37.860 --> 00:32:46.110
Philip Manlick: Oh, sorry, this is Phil, I was just gonna ask if we could shuffle things around, right. So I think like these are reports that UNAM when we always give them the people I'm thinking about an outside user

205
00:32:46.350 --> 00:32:53.910
Philip Manlick: It'll say, you know, D 13 see percent CD 15 and percent and then our carbon and nitrogen ratio. And if somebody wants to just copy and paste

206
00:32:55.020 --> 00:33:03.030
Philip Manlick: That chunk of their report and put it in. Could they reshuffle these headers to make them the order of the spreadsheet that they already

207
00:33:04.920 --> 00:33:23.880
Christopher Jordan: Know that's that's what that's what David is saying this order is important. There are, there can be variations within the groups but but if you if you significantly change the order, you're, you're definitely going to a point where it's no longer necessarily guaranteed to work.

208
00:33:24.600 --> 00:33:39.060
David Walling: Yeah, if we could technically make it work in any column order, we can handle all the reordering on the fly, but I keep advocating to the tech folks that I think trying to keep these things as consistent as possible is going to be a good thing.

209
00:33:40.620 --> 00:33:52.110
Mariel Lee Campbell: I think that having them consistent as possible on the doubt template download is really helpful to users. However, I see a problem with multiple users or even one user trying to

210
00:33:52.530 --> 00:34:04.500
Mariel Lee Campbell: Edit a file as most people are used to being able to shuffle their column headers around that that doesn't really matter. And if you want to add information later and having to constantly go back and follow some we would need really clear.

211
00:34:05.580 --> 00:34:14.370
Mariel Lee Campbell: Some clear interface instructions as to what the appropriate order should be to avoid problems down the line when people are actually working with the spreadsheets.

212
00:34:14.850 --> 00:34:20.640
Philip Manlick: I agree with Larry. I think that's pretty common that people are going to want to move these things around and at least putting it on the page.

213
00:34:20.670 --> 00:34:26.220
Philip Manlick: You know when you're creating your template that it needs to keep in that order, so people know because I think that a logical things people might

214
00:34:26.220 --> 00:34:34.230
Philip Manlick: Do is reshuffle these headers to match this spreadsheets that they already have, because it's easier to do that then data but you know spreadsheet that's already full of data.

215
00:34:34.560 --> 00:34:37.470
Mariel Lee Campbell: Yeah, cuz I mean just doing the test file.

216
00:34:38.820 --> 00:34:42.210
Mariel Lee Campbell: From the last time around I realized, Oh, well I could add a publication.

217
00:34:43.350 --> 00:34:52.770
Mariel Lee Campbell: In July, and do this. And oh, okay, I'm just gonna, I'm just gonna grab these two columns. I'm going to insert some new information and I just didn't put it wherever it felt like it made sense but I

218
00:34:53.640 --> 00:35:03.990
Mariel Lee Campbell: Know, I would want to have the flexibility that most people working with spreadsheets are used to that kind of flexibility and I think if we didn't have it there would need to be very, very explicit instructions and it might create some user difficulties.

219
00:35:05.370 --> 00:35:05.670
Mariel Lee Campbell: Okay.

220
00:35:05.700 --> 00:35:06.570
David Walling: Well, I

221
00:35:08.010 --> 00:35:18.900
David Walling: I would propose this let's make the ingest column order agnostic, but I would, however, when you hit this download button.

222
00:35:19.860 --> 00:35:31.560
David Walling: I would have it actually store the CSV file in a particular order, so it can interested in any order. But when it's actually saved and insert a backup. It should be in a more consistent folder semi

223
00:35:31.620 --> 00:35:33.000
Mariel Lee Campbell: I'd be great. Yeah, that's fine.

224
00:35:33.330 --> 00:35:34.920
Philip Manlick: That would be great. Yeah, it'd be awesome.

225
00:35:35.550 --> 00:35:38.940
Christopher Jordan: I think that that's technically not teaching so

226
00:35:39.240 --> 00:35:40.170
David Walling: I will explore that.

227
00:35:44.760 --> 00:35:54.930
David Walling: Right. So we took this last time as well. I think that we have this. So in addition to this kind of forum where you're actually selecting which field you're actually, once we also have a

228
00:35:54.960 --> 00:35:56.790
David Walling: More in depth field guide which it

229
00:35:56.790 --> 00:36:07.890
David Walling: opened up a new tab here it's going to show you your core analysis information. And for each of the metadata fields or core analysis information we

230
00:36:08.580 --> 00:36:24.060
David Walling: Provide the column name, whether it's required a little text about, you know what, what is this. So for lab, for example, it's the abbreviation for lab where the analysis was run and unique to lab, is that a matching laboratory must already exist in isolation.

231
00:36:25.710 --> 00:36:36.810
David Walling: We want lab to be like a first class entity with an ISO bank and so as we're starting to test and we're getting more labs in that lab likely may not already exist an ISO fake so

232
00:36:37.350 --> 00:36:49.410
David Walling: We want to talk about how to we don't want any lab, just to be entered by anybody right we want some kind of control over that. And we're still debating how we should handle that. So for now, you can just send a

233
00:36:50.700 --> 00:37:07.980
David Walling: An email to the ISO bake attack email list. If you get an error saying this lab doesn't exist just send us the information about the lab on it as an abbreviation that it's commonly known by give us that, and we'll go enter in through the admin form and have that available for your ingest

234
00:37:10.050 --> 00:37:13.170
David Walling: Sorry, that the field guide lab sample ID.

235
00:37:14.220 --> 00:37:16.800
David Walling: Again, it's required some of these fields are going to have

236
00:37:18.180 --> 00:37:29.070
David Walling: reg ex or other validators on the lab sample ID. It's free text that we want you to keep it between five and 50 characters and then only alpha numeric values with no spaces.

237
00:37:29.790 --> 00:37:40.170
David Walling: So we also provide some examples of what that looks like. Same for user sample ID same type of requirements on the field. And what's a valid value for that field.

238
00:37:41.640 --> 00:37:50.400
David Walling: What that field actually represents, and this is the sample this simple measurement measurement ID is the field that Chris pointed out at the beginning of the meeting.

239
00:37:51.210 --> 00:38:06.000
David Walling: That is our kind of new unique ID that essentially matches and uniquely identifies a line in a CSV file across actually all of ISO think right now, not just a given.

240
00:38:07.140 --> 00:38:18.000
David Walling: So it's five Charmin 50 char max alphanumeric no spaces in it. Again, this key must be unique for every record submitted to isolate

241
00:38:19.230 --> 00:38:25.140
David Walling: And you, there's the user can kind of come up with your own kind of unique formula for that. So, for instance, for this one.

242
00:38:25.980 --> 00:38:32.340
David Walling: We could just easily and this is easy to do an Excel or anything else just tack on a one or two, whatever.

243
00:38:33.120 --> 00:38:49.260
David Walling: If there's duplicate record say in your user sample it to easily create this unique sample mission and ID. Once we get to the point of updating records and as opposed to deleting and recreating will be king off of this particular field.

244
00:38:50.280 --> 00:38:51.360
To make those updates.

245
00:38:53.880 --> 00:38:58.710
David Walling: It's pretty important point. So any questions about that or confusion about what that represents

246
00:39:05.640 --> 00:39:08.580
Mariel Lee Campbell: Haven't you have a quick question question, David, is it

247
00:39:10.380 --> 00:39:19.140
Mariel Lee Campbell: Some of these ideas are going to differ from, for example, lab sample idea. You said it's alphanumeric no spaces, the data said I was provided actually

248
00:39:20.250 --> 00:39:25.830
Mariel Lee Campbell: Did have spaces and so you'll be changing the original sample ideas, it was

249
00:39:26.850 --> 00:39:32.400
Mariel Lee Campbell: Run in lab but course. Ideally, in the future, people will have some understanding of the standardization here.

250
00:39:33.690 --> 00:39:42.990
Mariel Lee Campbell: The same thing with sample measurement ID, is there a possibility that you would auto generate some of these values rather than putting the recording the user to do so to make it stand right

251
00:39:43.380 --> 00:39:52.740
David Walling: we've kind of gone back and forth on this. We talked about this using the primary key in the analysis table for this. Um, what

252
00:39:53.040 --> 00:39:58.680
David Walling: I guess we haven't come to a complete consensus, this isn't working for us in the past few weeks, and I guess it's

253
00:40:00.300 --> 00:40:03.300
David Walling: We're still considering open for debate. Chris, how we want to have on this.

254
00:40:04.140 --> 00:40:14.550
Christopher Jordan: Yeah i mean the the fundamental issue. You know, I thought about uniquely about, you know, just auto generating these in cases where

255
00:40:15.330 --> 00:40:33.120
Christopher Jordan: Where they're not supplied, but the, the issue is that this is an important field that is that is used to uniquely identify things so we don't at the point of ingest we don't have enough information to say that auto generating it is the correct thing to do.

256
00:40:33.900 --> 00:40:34.980
David Walling: A good point, actually.

257
00:40:35.550 --> 00:40:46.950
Christopher Jordan: So, Linda provide us to measurements right for the same sample and just omit a measurement ID and one and then not omit it in the next one.

258
00:40:47.490 --> 00:40:53.970
Christopher Jordan: And we could do the wrong thing. Basically, in process and parsing through the file. So, for the time being.

259
00:40:54.270 --> 00:41:01.050
Christopher Jordan: We're making this a required field and you'll see what what David has provided here is an example where you're basically just copying.

260
00:41:01.440 --> 00:41:11.520
Christopher Jordan: You're just copying a user sample ID into the sample measurement ID. And that's, that's just to tell the system. Yes, this is it. This is a unique measurement

261
00:41:11.910 --> 00:41:26.460
Christopher Jordan: That is associated with this particular sample, you know, we can talk in the future about if this is a case that we see a lot. Maybe there's a way to tell the ingest system. No, absolutely. I know what I'm doing. I want you to generate these for me.

262
00:41:27.330 --> 00:41:41.010
Christopher Jordan: But, but that's about the best we could do. We're never going to have the information to be confident that it's okay for us to you to generate those for you unless you explicitly tell us that that's the case.

263
00:41:41.550 --> 00:41:49.140
Mariel Lee Campbell: So just looking at your neon numbers right there. For example, having worked with neon and having just looked at them to a different researchers samples.

264
00:41:50.250 --> 00:41:51.660
Mariel Lee Campbell: Neon is pretty good about their

265
00:41:52.920 --> 00:42:02.970
Mariel Lee Campbell: Their formatting and it's pretty consistent in their way they do things but independent researchers are not the file. I was just working with head. Sometimes, there was a stat.

266
00:42:03.480 --> 00:42:19.800
Mariel Lee Campbell: Extra Space between a number and a letter and sometimes there wasn't. And it was and I could see problems with user error, causing a lot of conflicting values, we have the same issue in museum collections with coming up with global unique identifiers and I can see that.

267
00:42:20.820 --> 00:42:29.580
Mariel Lee Campbell: Kind of being an issue here that there will be error people submitting data will give you very multiple versions of the same thing. I don't know how much of it, you can control.

268
00:42:30.600 --> 00:42:42.450
Mariel Lee Campbell: But having one standard key that was assigned uniquely within is a bank that was outside of the control of users to create their own error might not be a bad idea, at least for the critical fields.

269
00:42:45.120 --> 00:42:55.530
Christopher Jordan: Yeah, and we will be clear, will we will have that you see on this, the far left hand side, this ID that is a unique to ISO bank ID that we generate

270
00:42:56.790 --> 00:43:15.780
Christopher Jordan: That that that uniquely identifies the values within ice. So we do have that we just need these fields so that so that people can, for example, give us multiple measurements that relate to the same sample and identify that that's

271
00:43:17.130 --> 00:43:17.820
Mariel Lee Campbell: What about

272
00:43:18.870 --> 00:43:21.240
Mariel Lee Campbell: Later correction or

273
00:43:22.740 --> 00:43:32.460
Mariel Lee Campbell: To any of these the same problem is a gen bake has the same problem. If a submitter puts in, for example, the wrong about your catalog number

274
00:43:34.500 --> 00:43:37.560
Mariel Lee Campbell: Or it's in the wrong format or it's just blatantly wrong.

275
00:43:38.910 --> 00:43:43.290
Mariel Lee Campbell: Then, and it's associating then that sample with the wrong museum voucher specimen.

276
00:43:43.800 --> 00:44:01.110
Mariel Lee Campbell: As a result, then it's only the submitter can make those corrections usually don't have the ability for museums to make corrections later. But there are needs. There have been issues where we do need to make corrections to associate evaluate the right samples. So I've just kind of

277
00:44:02.970 --> 00:44:06.330
Mariel Lee Campbell: From, from my perspective, working with Jim bank and with museums both

278
00:44:07.380 --> 00:44:17.160
Mariel Lee Campbell: Just heads up that there will be there will be problems down the line. And the more you is you know you have unique key on something. It just helps resolve those issues.

279
00:44:17.940 --> 00:44:18.150
Yeah.

280
00:44:25.980 --> 00:44:33.450
David Walling: Second session section is the core measurement information right now. That really means measurement type and then everything else is mostly

281
00:44:36.000 --> 00:44:47.820
David Walling: The analysis and measurement so meta data fields grouped into their categories for each of these. If you look at if you click the Help button, it'll pull up a

282
00:44:48.210 --> 00:44:49.260
David Walling: Additional page.

283
00:44:49.290 --> 00:44:56.220
David Walling: With a little easier to read formatting and this is where we can provide this more detailed description that could be quite

284
00:44:57.180 --> 00:45:09.210
David Walling: verbose. If it needed to be for Turkey Phil and additionally for the valid values could also have their own descriptions for each of those valid.us as I think Chris Thomas mentioned earlier.

285
00:45:10.470 --> 00:45:21.960
David Walling: And again, this is where we would be looking to the community to provide us or actually enter and help us manage this type of metadata information about the metadata.

286
00:45:27.060 --> 00:45:28.080
David Walling: I believe

287
00:45:35.910 --> 00:45:42.060
David Walling: Yeah. So as we turn you loose you more than likely in both with the upload templates.

288
00:45:44.010 --> 00:45:54.540
David Walling: Or actually trying to submit filled out templates run into errors. And so for now we're just going to ask you to provide again to the

289
00:45:55.500 --> 00:46:04.260
David Walling: Lists that I saw big attack that you've texted a to just shoot us an email with as much information about that area as you can so we can go track down exactly what's going on.

290
00:46:05.220 --> 00:46:11.430
David Walling: And just all use that as our primary feedback mechanism right now. Eventually, maybe we'll have some more kind of formal

291
00:46:11.850 --> 00:46:26.010
David Walling: Log a bug report type functionality, but as our community is still relatively small, we'll just keep a little more informal and flexible in that regard and Thomas Otter curses anything that I've missed that we wanted to cover for this particular

292
00:46:26.640 --> 00:46:38.220
Christopher Jordan: Will just on that, on that point. If you are if you're reporting something that you think is a is a bug that is you're getting an error on ingest and you believe you know

293
00:46:38.850 --> 00:46:47.790
Christopher Jordan: This shouldn't be giving me an error, then yes, please do report it to us. But ideally send us send us your CSV file as well.

294
00:46:47.820 --> 00:46:48.240
David Walling: Yes.

295
00:46:48.450 --> 00:47:01.050
Christopher Jordan: So that we can replicate the bug and and verify that it's fixed if it is and and I definitely do encourage you to do that. Try this stuff out, it will break. Don't worry about that will definitely break

296
00:47:01.830 --> 00:47:14.730
Christopher Jordan: But that's fine. That's, that's why we're doing this just, just let us know how you, how you think it's breaking incorrectly and give us the the CSV, so that we can validate the behavior when it's fixed

297
00:47:29.550 --> 00:47:35.160
Christopher Jordan: Okay. Any other, any other questions or comments on on any of that.

298
00:47:39.870 --> 00:47:48.420
Christopher Jordan: So yeah, as, as David said, we have a kind of a couple of other things to do just basically on on hardening the server. So we'll

299
00:47:48.750 --> 00:47:58.500
Christopher Jordan: We'll send you a follow up message within the next few days to let you know it's it's good to go. I won't necessarily discourage you from playing around with it as it as it is.

300
00:47:59.190 --> 00:48:10.290
Christopher Jordan: But, but just be aware that things, things may be changing quickly and coming up and down over the next couple of days. But certainly, you know, feel free to play around with things.

301
00:48:11.190 --> 00:48:15.450
Christopher Jordan: Look at the Field Guide. Again, I'll reiterate, if you see

302
00:48:15.930 --> 00:48:30.330
Christopher Jordan: A description there that you know we put in a minimal description and you think I have a much better description in this field, type it up and send it to us where we're happy to take all of the feedback we can get on that for the, for the time being.

303
00:48:35.430 --> 00:48:44.580
Christopher Jordan: If there's nothing else, then we can let you all go I did record the session. So I will, I will send out links to all of that for those of us who couldn't make it today.

304
00:48:50.430 --> 00:48:51.480
Christopher Jordan: Alright, thanks everyone.

305
00:48:54.000 --> 00:48:54.630
GABRIEL BOWEN: Thank you guys.

